Embedded Machine Learning · PhD course
Session 01
What is scarce on a small device, why it is still worth running learned models there, and the five numbers every later session prices its techniques against.
The framing
A model that is 94 % accurate and does not fit is 0 % useful.
Feasibility here is a hard constraint set, not a soft preference. The interesting content of the field is how accuracy trades against memory, latency and energy along that boundary.
How this course works
If ML is new to you: read the primer sections before each lecture. If hardware is new to you: the same.
The map
Method
Counting is not bookkeeping. Counting is the analysis.
01
What is actually inside the device — seen by someone who wants to run a model on it.
Anatomy
Five features
Orders of magnitude
A laptop has ~10⁴× the memory and ~10³× the power. Nothing transfers down without checking.
01
A hardware fact from 2004 decides the syllabus of a machine-learning course in 2026.
Foundations
Dennard scaling: shrinking a transistor also let you shrink V and C, so each generation ran faster at the same power. It ended in the mid-2000s, because below roughly 1 V the threshold voltage cannot fall further without leakage exploding.
Foundations
Everything after 2004 is parallelism and specialisation — because the free speed stopped arriving.
Definition by constraint
A rack-mounted automotive ECU is embedded. A Raspberry Pi on your desk mostly is not — and it is smaller.
The landscape
“Small” is never absolute. It is relative to one point on this plane.
02
A primer for those who have not trained models — and a reminder of what matters on a device for those who have.
Primer
The division of labour
Training happens on a workstation. Inference happens on the device.
When this course says "embedded ML" it almost always means embedded inference. Learning on the device is a research frontier (Session 4).
The workflow
Why on the device
These argue for putting the decision on the device — not necessarily the whole model.
03
Seven decades of data rate, eight of compute.
Where it is deployed
| Domain | Sensors | Example tasks | Usually binds |
|---|---|---|---|
| Audio and voice | MEMS microphone | keyword spotting, voice activity, acoustic events | always-on energy |
| Wearables and health | IMU, PPG, ECG | activity, falls, arrhythmia | battery, validation |
| Industrial monitoring | accelerometer, current | predictive maintenance, anomalies | scarce fault labels |
| Vision at the edge | low-res camera | person detection, counting | SRAM for activations |
| Environment, agriculture | gas, humidity, acoustic | air quality, pests | harvested energy |
| Automotive, robotics | radar, IMU, camera | driver monitoring, gestures | hard deadlines, safety |
Workloads
Sensor fusion
Break
Five minutes.
Next: the five budgets, and where the energy goes.
03
Given a device and an application, write these down before you open a notebook.
Budgets
The budget people get wrong
Parameters live at the back of a CNN; activations peak at the front, where the spatial resolution is still high. That asymmetry is the single most important structural fact about putting CNNs on microcontrollers.
Say it once, clearly
Peak activation memory, not parameter count, decides whether a CNN fits on a microcontroller.
Parameters at the back, activations at the front. A model can be tiny and still not run.
Worked derivation
The application wants 1 Hz. You are short by a factor of about 90.
The factor of 90
04
Students arrive believing arithmetic is expensive. It is not, and the gap is two orders of magnitude.
Foundations · Horowitz, ISSCC 2014 · 45 nm
| Operation | Energy | Operation | Energy |
|---|---|---|---|
| int8 add | 0.03 pJ | 32 kB cache read | 20 pJ |
| int32 add | 0.1 pJ | 1 MB cache read | 100 pJ |
| int8 multiply | 0.2 pJ | 8 kB cache read | 10 pJ |
| fp16 multiply | 1 pJ | DRAM access | 1.3–2.6 nJ |
| fp32 multiply | 4 pJ | ratio, DRAM : int8 add | ≈ 50 000× |
Absolute values shrink with process node. The ratios have proved remarkably stable.
The cost hierarchy
Radio is worse than DRAM. This bar chart is the quantitative case for edge inference.
Three conclusions
A “50 % FLOP reduction” that doubles memory traffic is a regression.
05
The architectural pattern that falls straight out of the energy picture.
Expected energy
04
The budgets only become real when applied.
Case studies
Keyword spotting on a coin cell
Vibration anomaly on a motor
Arrhythmia in a wearable
Exercise · 4 minutes, in pairs
There is rarely one right answer; there is always a wrong one: not writing them down.
Context
Signal processing on microcontrollers is decades old. What changed is which functions we are willing to learn rather than design.
Where a PhD fits
Before Session 2
V. Sze, Y.-H. Chen, T.-J. Yang, J. Emer, “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Proceedings of the IEEE 105(12), 2017 — §I–III. Optional: P. Warden, “Speech Commands,” 2018.
Why: The canonical statement that efficiency is a hardware–algorithm co-design problem, and the vocabulary — dataflow, reuse, energy per access — of Session 4.
Shortlist two presentation papers from the website; send your choice by the end of Session 2.
If you remember one thing
Write the five budgets down before you write any code. Arithmetic is nearly free; data movement and communication are not. When the numbers do not close, the four legitimate moves are cheaper inference, fewer inferences, more energy, or a different specification — never “train a better model”.
Before you go