Embedded Machine Learning · PhD course
Session 04
What the processor actually does with a network, the landscape from microcontrollers to neuromorphic chips, and the path from a trained graph to a measured, monitored binary.
The organising question
Which reuse does this machine’s dataflow capture — and does my layer have that kind of reuse?
That question predicts performance far better than peak TOPS. Keep it in your head for the whole session.
The thesis
A fast matrix multiply is not a cleverer algorithm. It is the same 2mnk multiply–accumulates, reorganised so each operand is loaded once and used many times.
What an embedded core can still offer
Out-of-order issue, speculation and deep caches are absent on a Cortex-M0+ or M4.
01
Big and fast is not something you buy. It is something you arrange.
Two machines, one principle
On a CPU the hardware decides and you hint. On an MCU you decide, in the linker script.
The trade, plotted
Tiling exists to keep you at the bottom left of this line.
Foundations
A microcontroller without a cache does not escape this; it inherits a harsher version, because no hardware guesses for you. The compensation is predictability: cycle counts on paper actually match the device.
02
A bound, not a prediction. It tells you what is impossible and how much headroom is left.
Derivation
π is peak compute, β is peak bandwidth. The ridge point is the arithmetic intensity a kernel must exceed to have any chance of reaching peak compute.
Three devices, one plot
The answer
| Layer | MACs | Bytes moved | Arithmetic intensity |
|---|---|---|---|
| Standard 3×3, 256→256 | 1.16 × 10⁸ | 6.9 × 10⁵ | ≈ 168 MAC/B |
| Depthwise 3×3, 256 | 4.5 × 10⁵ | 1.0 × 10⁵ | ≈ 4.4 MAC/B |
| Pointwise 1×1, 256→256 | 1.28 × 10⁷ | 1.7 × 10⁵ | ≈ 77 MAC/B |
| Dense, 1024→1024 | 1.05 × 10⁶ | 1.05 × 10⁶ | ≈ 1 MAC/B |
AI = MACs / (bytes in + bytes of weights + bytes out), at batch 1. Depthwise and dense read almost as many bytes as they perform multiplies.
Two worked ridge points
A genuinely useful asymmetry: on an MCU, reducing MACs often really does reduce time.
03
From the naive triple loop to a production kernel.
Step 2, made visible
Grey rectangles are cache lines. Copper cells are the words the loop actually uses.
Why blocking works
Arithmetic intensity scales with b — and b is capped by the fast memory that must hold three tiles.
Convolution as GEMM
Convolution on a CPU
| Strategy | Arithmetic | Extra memory | Note |
|---|---|---|---|
| im2col + GEMM | baseline | ×k² patch buffer | Reuses a tuned GEMM; the blow-up is often fatal on an MCU |
| Partial im2col | baseline | two columns | CMSIS-NN’s approach — bounded SRAM |
| Implicit GEMM | baseline | none | Index arithmetic replaces the copy; standard on GPUs |
| Direct convolution | baseline | none | Best for depthwise |
| Winograd F(2×2,3×3) | 16 vs 36 mults → 2.25× | transform buffers | Numerically fragile in low precision |
| FFT convolution | wins for large k | complex buffers | Rarely worth it for 3×3; relevant for long 1-D audio kernels |
Layout matters as much as algorithm: NHWC keeps channels innermost and contiguous, which is what a vectorised inner product over channels wants.
01
What each accelerator accelerates — and what it demands of the model.
Four features that decide your design
Matrix multiply here is Session 4 with the constants changed: no cache to block for, partial im2col, int8 weights with int32 accumulators, SMLAD inner loops.
Case study
The third core in the room
Modern audio SoCs pair all three, and the interesting design question is which stage runs where.
Three mechanisms, three consequences
GEMM on a GPU is Session 4’s five steps with the fast memory made explicit: shared memory is software-managed.
The cost of a branch
Why unstructured sparsity, early exits and any data-dependent branch are expensive here — and why N:M sparsity, fixed at compile time, is not.
Break
Five minutes.
Next: dataflow accelerators, beyond-digital hardware, and deployment.
Trace it cycle by cycle
Google’s first-generation TPU: a 256×256 array of 8-bit MAC units — 65 536 multipliers, reported at a peak of 92 TOPS.
The taxonomy that lets you read a datasheet
A depthwise layer has almost no weight reuse to capture. Weight-stationary hardware will not save it.
At the microcontroller end
Check the operator-support list before designing the architecture, not after.
Two physical laws instead of a datapath
A 64-core PCM chip in 14 nm reported up to 63.1 TOPS and 9.76 TOPS/W in low-precision mode, with near-software-equivalent accuracy on ResNet and LSTM workloads (Le Gallo et al., Nature Electronics, 2023).
Honestly
Present it as a live research direction with a real physical argument — not as a product category.
Choosing
Putting it together
| If your binding constraint is… | Consider | Watch out for |
|---|---|---|
| Microwatts, years on a cell | MCU, duty-cycled, cascade front end | Wake-up energy dominating compute |
| Milliwatts, int8 CNN, tight unit cost | MCU + microNPU | Unsupported operators falling back to CPU |
| Single-stream latency | Mobile SoC / embedded module | Batch-1 occupancy collapse on the GPU |
| Many concurrent streams | Embedded GPU module | Thermal throttling changing your benchmark |
| Deterministic microsecond latency | FPGA | Development cost and model rigidity |
| Volume > 10⁶, fixed model | ASIC | Non-recurring engineering and the schedule |
In practice ecosystem maturity — compiler quality, operator coverage, debugging tools — decides more projects than peak performance, and it is the factor most often omitted from comparisons.
02
Every stage can break the model silently.
Seven stages, seven silent failures
Parity-test at every arrow: same input, both sides, compare per tensor. The first arrow where SQNR collapses is the bug.
Offline packing
A long residual connection raises the arena for the whole network — which is why residual topology is a memory decision.
The observation the whole paper rests on
Only a handful of early layers do not fit — and the whole network is sized for them.
Methodology for everyone
A paper reporting “11.8 ms” has reported the left edge of this picture.
How to measure on an MCU
DWT->CYCCNT) around the kernel and subtract the measurement overhead.A single run on a system with interrupts enabled is not a measurement.
Testing
Budgets that are not enforced are not budgets.
Name it before you fix it
Covariate → recalibrate. Concept → needs new labels. Label shift → correctable in closed form if the model is calibrated.
The quantitative reason
Backpropagation must retain every forward activation until its gradient is used. Every practical method is a way of storing fewer of them.
Threat model
| Threat | What the attacker gets | Mitigation |
|---|---|---|
| Model extraction | the weights, by reading flash or by querying | Secure boot, encrypted flash, read-out protection, rate limiting |
| Membership inference | whether a record was in the training set | Differential privacy; avoid over-fitting |
| Adversarial inputs | controlled misclassification | Adversarial training, input plausibility checks |
| Physical / sensor spoofing | injected signals treated as real — ultrasonic audio, projected light | Multi-sensor consistency, band-limiting in the analog chain |
| Power / EM side channel | architecture and sometimes weights, from the current trace | Constant-time kernels, masking, noise injection — all costly |
| Fault injection | skipped instructions, corrupted results | Redundant computation, output plausibility checks |
The side-channel row is the one specific to this domain: physical access means the current trace is correlated with the computation.
Frontier
Language models at the edge
Real
Not real
Session 5
One landmark paper per presenter, one topic area each (A quantization · B pruning and distillation · C efficient architectures and NAS · D TinyML systems and hardware). 15 minutes + 5 minutes discussion.
Why: Claim in one sentence with its scope; the mechanism re-derivable; the load-bearing experiment; one real limitation; what it means on a constrained device.
Slides to the lecturer the evening before. Discussants: one prepared question.
If you remember one thing
Ask where the bytes come from before asking how fast the arithmetic is. The roofline says which layers an accelerator can help, the toolchain says which it can run at all, and only a measurement on the device, with its method stated, says what you actually have.
Before you go