EmbML · S1 · Embedded systems1 / 44

Embedded Machine Learning · PhD course

Session 01

Embedded systems and the case for machine learning at the edge

What is scarce on a small device, why it is still worth running learned models there, and the five numbers every later session prices its techniques against.

90 minutesSession 1 of 5
EmbML · S1 · Embedded systems2 / 44

The framing

A model that is 94 % accurate and does not fit is 0 % useful.

Feasibility here is a hard constraint set, not a soft preference. The interesting content of the field is how accuracy trades against memory, latency and energy along that boundary.

EmbML · S1 · Embedded systems3 / 44

How this course works

Four lectures, one presentation session, two layers of material

  • Sessions 1–4: the pipeline from sensor to silicon — the device and its budgets, the data, the models, the hardware.
  • Session 5: three or four of you present one landmark paper each, from four topic areas.
  • Web chapters go further than the lecture: primer sections for missing background, depth sections for the full argument.
  • Choose your presentation paper by the end of Session 2.

If ML is new to you: read the primer sections before each lecture. If hardware is new to you: the same.

EmbML · S1 · Embedded systems4 / 44

The map

The course follows the pipeline

EmbML · S1 · Embedded systems5 / 44

Method

Four questions we ask of every technique

  • What does it cost? Count MACs, weight bytes, peak activation bytes, and bytes moved.
  • Where does the cost go? Arithmetic is nearly free; data movement is not.
  • What does it cost in accuracy — and why? Mechanisms, not folklore.
  • How would you know? Which device, compiler, clock, how many repetitions, what dispersion?

Counting is not bookkeeping. Counting is the analysis.

EmbML · S1 · Embedded systems6 / 44

01

Embedded systems, in plain terms

What is actually inside the device — seen by someone who wants to run a model on it.

EmbML · S1 · Embedded systems7 / 44

Anatomy

A microcontroller, as an ML engineer sees it

EmbML · S1 · Embedded systems8 / 44

Five features

What shapes everything that follows

  • Two small memories: flash (weights, code) and SRAM (activations, buffers) — kB to a few MB.
  • No operating system, or a tiny real-time one: no virtual memory, often no heap.
  • A modest core: tens to hundreds of MHz, with SIMD that does two 16-bit MACs per instruction.
  • Peripherals and DMA: samples arrive while the core sleeps.
  • Power modes: average power is set by how rarely and how briefly it wakes.
EmbML · S1 · Embedded systems9 / 44

Orders of magnitude

Numbers to carry around

Flash0.25–4 MBweights and program
SRAM32 kB–3 MBactivations, buffers, stack
Clock48–480 MHzCortex-M class core
PowerµW–100 mWsleep to full speed

A laptop has ~10⁴× the memory and ~10³× the power. Nothing transfers down without checking.

EmbML · S1 · Embedded systems10 / 44

01

Why this field exists

A hardware fact from 2004 decides the syllabus of a machine-learning course in 2026.

EmbML · S1 · Embedded systems11 / 44

Foundations

Where the constraint comes from

P_dyn = α · C · V² · f dynamic: paid on every switch
P_stat = leakage paid whether you switch or not

Dennard scaling: shrinking a transistor also let you shrink V and C, so each generation ran faster at the same power. It ended in the mid-2000s, because below roughly 1 V the threshold voltage cannot fall further without leakage exploding.

EmbML · S1 · Embedded systems12 / 44

Foundations

Transistors kept doubling. Clock speed did not.

Everything after 2004 is parallelism and specialisation — because the free speed stopped arriving.

EmbML · S1 · Embedded systems13 / 44

Definition by constraint

Three dimensions that actually matter

  • Timing. Hard real-time fails on a missed deadline; soft real-time degrades. Most ML sits soft inside a hard host — and that interface is where the bugs live.
  • Resources. Kilobytes to a few megabytes. No virtual memory. Often no heap.
  • Energy. A fixed store and a required service life — which converts directly into a power ceiling.

A rack-mounted automotive ECU is embedded. A Raspberry Pi on your desk mostly is not — and it is smaller.

EmbML · S1 · Embedded systems14 / 44

The landscape

Six orders of magnitude of power, seven of memory

“Small” is never absolute. It is relative to one point on this plane.

EmbML · S1 · Embedded systems15 / 44

02

What machine learning adds

A primer for those who have not trained models — and a reminder of what matters on a device for those who have.

EmbML · S1 · Embedded systems16 / 44

Primer

A model, in one slide

  • A model is a function ŷ = f(x; θ): input window → label, number or score.
  • Its form is chosen; its parameters θ are learned from labelled examples.
  • Training searches for θ — expensive, done on a workstation or in the cloud.
  • Inference evaluates f once for a new input — cheap, done on the device.
  • Generalisation is the goal: right answers on inputs never seen before.
EmbML · S1 · Embedded systems17 / 44

The division of labour

Training happens on a workstation. Inference happens on the device.

When this course says "embedded ML" it almost always means embedded inference. Learning on the device is a research frontier (Session 4).

EmbML · S1 · Embedded systems18 / 44

The workflow

Most of the effort is in the two loops

EmbML · S1 · Embedded systems19 / 44

Why on the device

Five reasons — in the order they are usually the real one

  • Energy of communication: sending raw data costs more than deciding locally.
  • Privacy and regulation: data that never leaves cannot leak.
  • Latency and its variance: control loops care about the tail, not the mean.
  • Availability: fields, mines, aircraft, oceans.
  • Cost at scale: bandwidth and cloud inference are recurring costs.

These argue for putting the decision on the device — not necessarily the whole model.

EmbML · S1 · Embedded systems20 / 44

03

Applications and workloads

Seven decades of data rate, eight of compute.

EmbML · S1 · Embedded systems21 / 44

Where it is deployed

Application domains and what usually binds

DomainSensorsExample tasksUsually binds
Audio and voiceMEMS microphonekeyword spotting, voice activity, acoustic eventsalways-on energy
Wearables and healthIMU, PPG, ECGactivity, falls, arrhythmiabattery, validation
Industrial monitoringaccelerometer, currentpredictive maintenance, anomaliesscarce fault labels
Vision at the edgelow-res cameraperson detection, countingSRAM for activations
Environment, agriculturegas, humidity, acousticair quality, pestsharvested energy
Automotive, roboticsradar, IMU, cameradriver monitoring, gestureshard deadlines, safety
EmbML · S1 · Embedded systems22 / 44

Workloads

Data rate decides whether streaming is possible; compute decides the device

EmbML · S1 · Embedded systems23 / 44

Sensor fusion

Where to merge several sensors

EmbML · S1 · Embedded systems24 / 44

Break

Five minutes.

Next: the five budgets, and where the energy goes.

EmbML · S1 · Embedded systems25 / 44

03

The five budgets

Given a device and an application, write these down before you open a notebook.

EmbML · S1 · Embedded systems26 / 44

Budgets

Five numbers, in the order they usually bind

  • Weight memory — flash. P·b/8 bytes plus metadata, interpreter and the rest of the firmware. The model gets about half the flash, never all of it.
  • Peak activation memory — SRAM. Not the model size. This is the one that kills naive designs.
  • Compute — MACs, divided by the device’s sustained rate, never its peak. A lower bound only.
  • Latency — set by the application, and it must include feature extraction.
  • Energy per inference — and the duty cycle that follows from it.
EmbML · S1 · Embedded systems27 / 44

The budget people get wrong

Peak activation memory

peak_SRAM = max over layers i of ( |x_i| + |y_i| + Σ |t| for t still live )
Conv2D MACs = H · W · k² · C_in · C_out

Parameters live at the back of a CNN; activations peak at the front, where the spatial resolution is still high. That asymmetry is the single most important structural fact about putting CNNs on microcontrollers.

EmbML · S1 · Embedded systems28 / 44

Say it once, clearly

Peak activation memory, not parameter count, decides whether a CNN fits on a microcontroller.

Parameters at the back, activations at the front. A model can be tiny and still not run.

EmbML · S1 · Embedded systems29 / 44

Worked derivation

A one-year coin cell

  1. A CR2032 is nominally 3 V with about 225 mAh: E = 0.225 × 3 × 3600 ≈ 2.4 × 10³ J
  2. One year is 3.15 × 10⁷ s, so the average power ceiling is P ≤ 2.4×10³ / 3.15×10⁷ ≈ 77 µW
  3. Standby (RTC, sensor bias, leakage) already takes 20 µW → 57 µW left for inference
  4. If one inference costs 5 mJ: f ≤ 57 µW / 5 mJ ≈ 0.011 Hz — one inference every ~90 s

The application wants 1 Hz. You are short by a factor of about 90.

EmbML · S1 · Embedded systems30 / 44

The factor of 90

Four legitimate ways out — and one that is not

  • Reduce energy per inference — compression and efficient architectures. Sessions 8–12.
  • Reduce the rate — cascade with a cheap always-on stage. §1.5, and in ten minutes.
  • Increase the store — a bigger cell, or energy harvesting.
  • Relax the specification — the most underused option in the industry.
  • Not on the list: “train a better model”.
EmbML · S1 · Embedded systems31 / 44

04

Where the energy actually goes

Students arrive believing arithmetic is expensive. It is not, and the gap is two orders of magnitude.

EmbML · S1 · Embedded systems32 / 44

Foundations · Horowitz, ISSCC 2014 · 45 nm

Energy per operation

OperationEnergyOperationEnergy
int8 add0.03 pJ32 kB cache read20 pJ
int32 add0.1 pJ1 MB cache read100 pJ
int8 multiply0.2 pJ8 kB cache read10 pJ
fp16 multiply1 pJDRAM access1.3–2.6 nJ
fp32 multiply4 pJratio, DRAM : int8 add≈ 50 000×

Absolute values shrink with process node. The ratios have proved remarkably stable.

EmbML · S1 · Embedded systems33 / 44

The cost hierarchy

Arithmetic on the left. Moving the same operands on the right.

Radio is worse than DRAM. This bar chart is the quantitative case for edge inference.

EmbML · S1 · Embedded systems34 / 44

Three conclusions

What follows from that picture

  • Precision is an energy knob. fp32 → int8 multiply is roughly 20× in the multiplier alone, before the halved memory traffic. Session 3.
  • Locality is a bigger knob than precision. Any technique that keeps operands on-chip beats any technique that merely reduces arithmetic. Sessions 6–7.
  • Sparsity only pays if the hardware can skip. Removing a multiply saves 0.2 pJ; the index bookkeeping may cost more. Session 3.

A “50 % FLOP reduction” that doubles memory traffic is a regression.

EmbML · S1 · Embedded systems35 / 44

05

The cascade

The architectural pattern that falls straight out of the energy picture.

EmbML · S1 · Embedded systems36 / 44

Expected energy

Why a 100× more expensive stage can be nearly free

  • E[total] = E₁ + p₁·E₂ + p₁p₂·E₃ + …
  • Most windows are negative: no keyword, no person, no fault.
  • With p₁ = 2 %, a stage 100× more expensive costs only 2× the cheap stage.
  • So the always-on stage is tuned for recall at fixed energy. Precision is the next stage’s job.
EmbML · S1 · Embedded systems37 / 44

04

Three systems, end to end

The budgets only become real when applied.

EmbML · S1 · Embedded systems38 / 44

Case studies

Three systems, three different binding constraints

Keyword spotting on a coin cell

  • 16 kHz mic → 49×10 MFCC map
  • DS-CNN: ≈39 k params, 2.7 M MACs
  • Cortex-M4F class MCU
  • Binds: energy — fixed by a voice-activity cascade

Vibration anomaly on a motor

  • 3-axis accelerometer, 1–10 kHz
  • band energies per second
  • autoencoder trained on healthy data
  • Binds: data — no labelled failures, drifting baseline

Arrhythmia in a wearable

  • single-lead ECG, 250 Hz
  • intervals + tree, or a 1-D CNN
  • wearable SoC
  • Binds: validation — sensitivity on clinical data
EmbML · S1 · Embedded systems39 / 44

Exercise · 4 minutes, in pairs

Find the binding budget

  1. Pick one of the three systems.
  2. Write the five budgets with an order of magnitude each: flash, peak SRAM, MACs, latency, energy per inference.
  3. Circle the one that binds.
  4. Change one thing — a camera instead of a microphone, mains power instead of a coin cell, 10× more data — and say which budget binds now.

There is rarely one right answer; there is always a wrong one: not writing them down.

EmbML · S1 · Embedded systems40 / 44

Context

Four inflection points, honestly told

  • 2016–17 — practical mobile-scale CNNs. Session 3.
  • 2017–18 — integer-only inference and a toolchain that made it routine. Session 3.
  • 2020 → — memory-aware co-design of architecture and runtime for MCUs. Session 4.
  • 2023 → — transformers, state-space models and language models at the edge. Open research, not settled practice. Session 4.

Signal processing on microcontrollers is decades old. What changed is which functions we are willing to learn rather than design.

EmbML · S1 · Embedded systems41 / 44

Where a PhD fits

Open problems, by session

  • On-device learning within kB of memory, without labels (S4).
  • Data efficiency: self-supervision, few-shot adaptation, synthetic data (S2–S3).
  • Hardware-aware design: NAS and compression on measured latency and energy (S3).
  • Compilers and runtimes for heterogeneous MCU + NPU systems (S4).
  • Beyond digital: analog in-memory and neuromorphic hardware (S4).
  • Trust: robustness, privacy, secure updates — and honest measurement.
EmbML · S1 · Embedded systems42 / 44

Before Session 2

Reading, and choosing a paper

V. Sze, Y.-H. Chen, T.-J. Yang, J. Emer, “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Proceedings of the IEEE 105(12), 2017 — §I–III. Optional: P. Warden, “Speech Commands,” 2018.

Why: The canonical statement that efficiency is a hardware–algorithm co-design problem, and the vocabulary — dataflow, reuse, energy per access — of Session 4.

Shortlist two presentation papers from the website; send your choice by the end of Session 2.

EmbML · S1 · Embedded systems43 / 44

If you remember one thing

Write the five budgets down before you write any code. Arithmetic is nearly free; data movement and communication are not. When the numbers do not close, the four legitimate moves are cheaper inference, fewer inferences, more energy, or a different specification — never “train a better model”.

EmbML · S1 · Embedded systems44 / 44

Before you go

Think about these

  1. The “privacy” argument for edge inference is often made loosely. Construct a case where moving inference to the device makes privacy worse.
  2. A vendor advertises “2 TOPS at 1 W”. Name four things you must know before that number predicts anything about your application.
  3. For your own research problem: which fusion strategy, and which budget binds first?
S1 · Embedded systems