EmbML · S4 · Hardware & deployment1 / 47

Embedded Machine Learning · PhD course

Session 04

Hardware platforms and deployment

What the processor actually does with a network, the landscape from microcontrollers to neuromorphic chips, and the path from a trained graph to a measured, monitored binary.

90 minutesSession 4 of 5
EmbML · S4 · Hardware & deployment2 / 47

The organising question

Which reuse does this machine’s dataflow capture — and does my layer have that kind of reuse?

That question predicts performance far better than peak TOPS. Keep it in your head for the whole session.

EmbML · S4 · Hardware & deployment3 / 47

The thesis

A fast matrix multiply is not a cleverer algorithm. It is the same 2mnk multiply–accumulates, reorganised so each operand is loaded once and used many times.

EmbML · S4 · Hardware & deployment4 / 47

What an embedded core can still offer

Pipelining and SIMD

Out-of-order issue, speculation and deep caches are absent on a Cortex-M0+ or M4.

EmbML · S4 · Hardware & deployment5 / 47

01

The memory hierarchy

Big and fast is not something you buy. It is something you arrange.

EmbML · S4 · Hardware & deployment6 / 47

Two machines, one principle

Who decides what is close?

On a CPU the hardware decides and you hint. On an MCU you decide, in the linker script.

EmbML · S4 · Hardware & deployment7 / 47

The trade, plotted

1000× more capacity costs about 1000× more latency

Tiling exists to keep you at the bottom left of this line.

EmbML · S4 · Hardware & deployment8 / 47

Foundations

Locality, in two rules

Memory is delivered in LINES (32 or 64 bytes), never in words.
Spatial locality → touch the rest of a line before it is evicted
Temporal locality → if you will need a value again, use it again soon

A microcontroller without a cache does not escape this; it inherits a harsher version, because no hardware guesses for you. The compensation is predictability: cycle counts on paper actually match the device.

EmbML · S4 · Hardware & deployment9 / 47

02

The roofline model

A bound, not a prediction. It tells you what is impossible and how much headroom is left.

EmbML · S4 · Hardware & deployment10 / 47

Derivation

Two lower bounds on time

I = W / Q arithmetic intensity (ops per byte)
T ≥ max( W/π , Q/β ) compute bound, memory bound
P = W/T ≤ min( π , β · I )
ridge point I* = π / β

π is peak compute, β is peak bandwidth. The ridge point is the arithmetic intensity a kernel must exceed to have any chance of reaching peak compute.

EmbML · S4 · Hardware & deployment11 / 47

Three devices, one plot

The same layer has a different bottleneck on each device

EmbML · S4 · Hardware & deployment12 / 47

The answer

Arithmetic intensity, computed for a 14×14×256 int8 activation

LayerMACsBytes movedArithmetic intensity
Standard 3×3, 256→2561.16 × 10⁸6.9 × 10⁵≈ 168 MAC/B
Depthwise 3×3, 2564.5 × 10⁵1.0 × 10⁵≈ 4.4 MAC/B
Pointwise 1×1, 256→2561.28 × 10⁷1.7 × 10⁵≈ 77 MAC/B
Dense, 1024→10241.05 × 10⁶1.05 × 10⁶≈ 1 MAC/B

AI = MACs / (bytes in + bytes of weights + bytes out), at batch 1. Depthwise and dense read almost as many bytes as they perform multiplies.

EmbML · S4 · Hardware & deployment13 / 47

Two worked ridge points

The bar is much lower on a microcontroller

Laptop CPU core5≈100 GFLOP/s over ≈20 GB/s → I* = 5 FLOP/byte
Cortex-M4F @ 80 MHz0.5≈160 MMAC/s over ≈320 MB/s → I* = 0.5 MAC/byte
Dense layer, batch 11≈1 MAC/byte — hopeless on the laptop, fine on the MCU
Depthwise 3×34.4MAC/byte — above the M4F ridge, below a GPU's

A genuinely useful asymmetry: on an MCU, reducing MACs often really does reduce time.

EmbML · S4 · Hardware & deployment14 / 47

03

GEMM in five steps

From the naive triple loop to a production kernel.

EmbML · S4 · Hardware & deployment15 / 47

Step 2, made visible

Same arithmetic, four times the traffic

Grey rectangles are cache lines. Copper cells are the words the loop actually uses.

EmbML · S4 · Hardware & deployment16 / 47

Why blocking works

Load 2b² elements, do b³ multiply–accumulates

Arithmetic intensity scales with b — and b is capped by the fast memory that must hold three tiles.

EmbML · S4 · Hardware & deployment17 / 47

Convolution as GEMM

im2col: reuse the fast kernel, pay in memory

EmbML · S4 · Hardware & deployment18 / 47

Convolution on a CPU

Six strategies, two of which change the numerics

StrategyArithmeticExtra memoryNote
im2col + GEMMbaseline×k² patch bufferReuses a tuned GEMM; the blow-up is often fatal on an MCU
Partial im2colbaselinetwo columnsCMSIS-NN’s approach — bounded SRAM
Implicit GEMMbaselinenoneIndex arithmetic replaces the copy; standard on GPUs
Direct convolutionbaselinenoneBest for depthwise
Winograd F(2×2,3×3)16 vs 36 mults → 2.25×transform buffersNumerically fragile in low precision
FFT convolutionwins for large kcomplex buffersRarely worth it for 3×3; relevant for long 1-D audio kernels

Layout matters as much as algorithm: NHWC keeps channels innermost and contiguous, which is what a vectorised inner product over channels wants.

EmbML · S4 · Hardware & deployment19 / 47

01

The platform landscape

What each accelerator accelerates — and what it demands of the model.

EmbML · S4 · Hardware & deployment20 / 47

Four features that decide your design

Planning an MCU inference

  • A flat, explicit memory map. No virtual memory, an MPU at most. You choose — in the linker script — what lives in flash, SRAM or TCM.
  • DMA. Acquire the next audio buffer while computing on the previous one. Double buffering is the difference between meeting a deadline and dropping frames.
  • Power states. Run, sleep, deep sleep, retention — with wake-up latencies from microseconds to milliseconds. Wake-up energy frequently dominates the compute.
  • Interrupt-driven acquisition. A short, bounded ISR; inference in the main loop. Get this wrong and you get sporadic sample loss that looks exactly like a model problem.

Matrix multiply here is Session 4 with the constants changed: no cache to block for, partial im2col, int8 weights with int32 accumulators, SMLAD inner loops.

EmbML · S4 · Hardware & deployment21 / 47

Case study

A keyword spotter on an Apollo4-class MCU, every assumption visible

EmbML · S4 · Hardware & deployment22 / 47

The third core in the room

Digital signal processors

  • Harvard architecture — separate instruction and data buses, so a MAC can fetch an operand and an instruction in the same cycle.
  • Single-cycle multiply–accumulate with wide guard-bit accumulators.
  • Hardware circular addressing — free ring buffers for FIR filters and streaming windows.
  • Zero-overhead loops, and often VLIW issue scheduled statically by the compiler.
  • For an always-on audio front end — filtering, FFT, mel filterbank — a DSP core is frequently more efficient than either the MCU or the NPU beside it.

Modern audio SoCs pair all three, and the interesting design question is which stage runs where.

EmbML · S4 · Hardware & deployment23 / 47

Three mechanisms, three consequences

What a GPU actually is

  • Threads are scheduled in lock-step groups — warps, typically 32. All lanes execute the same instruction.
  • Divergence. If lanes take different branches, both paths execute with inactive lanes masked.
  • Coalescing. Consecutive threads should touch consecutive addresses. Layout choices are coalescing choices.
  • Occupancy. Latency is hidden by resident warps; residency is limited by registers and shared memory. Tile sizing is a three-way negotiation.

GEMM on a GPU is Session 4’s five steps with the fast memory made explicit: shared memory is software-managed.

EmbML · S4 · Hardware & deployment24 / 47

The cost of a branch

Both sides run. Inactive lanes still occupy their slot in time.

Why unstructured sparsity, early exits and any data-dependent branch are expensive here — and why N:M sparsity, fixed at compile time, is not.

EmbML · S4 · Hardware & deployment25 / 47

Break

Five minutes.

Next: dataflow accelerators, beyond-digital hardware, and deployment.

EmbML · S4 · Hardware & deployment26 / 47

Trace it cycle by cycle

N² multiply–accumulates per cycle, 2N operands fetched

Google’s first-generation TPU: a 256×256 array of 8-bit MAC units — 65 536 multipliers, reported at a peak of 92 TOPS.

EmbML · S4 · Hardware & deployment27 / 47

The taxonomy that lets you read a datasheet

Which reuse does it capture?

A depthwise layer has almost no weight reuse to capture. Weight-stationary hardware will not save it.

EmbML · S4 · Hardware & deployment28 / 47

At the microcontroller end

MicroNPUs, and the operator-coverage trap

  • A small accelerator beside a Cortex-M core, executing a restricted operator set in int8 from a command stream prepared offline by a vendor compiler.
  • Anything unsupported falls back to the CPU — often with a layout conversion in each direction.
  • One unsupported operator in the middle of a graph can cost more than every optimisation elsewhere.
  • The vendor compiler, not your training framework, decides whether your model is fast.

Check the operator-support list before designing the architecture, not after.

EmbML · S4 · Hardware & deployment29 / 47

Two physical laws instead of a datapath

Analog in-memory computing

A 64-core PCM chip in 14 nm reported up to 63.1 TOPS and 9.76 TOPS/W in low-precision mode, with near-software-equivalent accuracy on ResNet and LSTM workloads (Le Gallo et al., Nature Electronics, 2023).

EmbML · S4 · Hardware & deployment30 / 47

Honestly

Neuromorphic and spiking systems

  • Information in the timing of sparse events; computation only when events arrive.
  • A natural fit for always-on sensing with a low event rate — event cameras, spiking audio front ends.
  • Real demonstrated advantages in latency and energy for sparse, temporally structured signals.
  • Open problems: training (surrogate gradients are improving, not routine), no standard software stack, and a shortage of tasks where the sparsity assumption clearly holds.

Present it as a live research direction with a real physical argument — not as a product category.

EmbML · S4 · Hardware & deployment31 / 47

Choosing

No platform wins every axis

EmbML · S4 · Hardware & deployment32 / 47

Putting it together

Choosing a platform

If your binding constraint is…ConsiderWatch out for
Microwatts, years on a cellMCU, duty-cycled, cascade front endWake-up energy dominating compute
Milliwatts, int8 CNN, tight unit costMCU + microNPUUnsupported operators falling back to CPU
Single-stream latencyMobile SoC / embedded moduleBatch-1 occupancy collapse on the GPU
Many concurrent streamsEmbedded GPU moduleThermal throttling changing your benchmark
Deterministic microsecond latencyFPGADevelopment cost and model rigidity
Volume > 10⁶, fixed modelASICNon-recurring engineering and the schedule

In practice ecosystem maturity — compiler quality, operator coverage, debugging tools — decides more projects than peak performance, and it is the factor most often omitted from comparisons.

EmbML · S4 · Hardware & deployment33 / 47

02

From graph to binary — and after

Every stage can break the model silently.

EmbML · S4 · Hardware & deployment34 / 47

Seven stages, seven silent failures

The deployment pipeline

Parity-test at every arrow: same input, both sides, compare per tensor. The first arrow where SQNR collapses is the bug.

EmbML · S4 · Hardware & deployment35 / 47

Offline packing

Memory arena planning

A long residual connection raises the arena for the whole network — which is why residual topology is a memory decision.

EmbML · S4 · Hardware & deployment36 / 47

The observation the whole paper rests on

Peak memory is a spike, not a plateau

Only a handful of early layers do not fit — and the whole network is sized for them.

EmbML · S4 · Hardware & deployment37 / 47

Methodology for everyone

A measurement is a distribution, not a number

A paper reporting “11.8 ms” has reported the left edge of this picture.

EmbML · S4 · Hardware & deployment38 / 47

How to measure on an MCU

A defensible latency measurement

  1. Read the cycle counter (DWT->CYCCNT) around the kernel and subtract the measurement overhead.
  2. Report cycles, not milliseconds — the number then survives a clock change.
  3. Decide and state whether caches are warm. Either is fine; silence is not.
  4. Repeat, and report the median and the interquartile range.
  5. Say whether interrupts were enabled.

A single run on a system with interrupts enabled is not a measurement.

EmbML · S4 · Hardware & deployment39 / 47

Testing

Four practices that pay for themselves in a week

  1. Per-layer parity. Identical inputs through the float reference and the deployed graph; dump every intermediate; report per-layer SQNR.
  2. Golden vectors. Freeze inputs and expected outputs into the firmware test suite, so a toolchain upgrade cannot silently change behaviour.
  3. Hardware in the loop. CI that flashes real boards. Emulators do not reproduce timing, DMA, interrupts or wait states.
  4. Resource assertions. Fail the build if flash, arena size or worst-case latency exceeds the budget from Session 1.

Budgets that are not enforced are not budgets.

EmbML · S4 · Hardware & deployment40 / 47

Name it before you fix it

Three drifts, three different remedies

Covariate → recalibrate. Concept → needs new labels. Label shift → correctable in closed form if the model is calibrated.

EmbML · S4 · Hardware & deployment41 / 47

The quantitative reason

Training memory is activations, not weights

Backpropagation must retain every forward activation until its gradient is used. Every practical method is a way of storing fewer of them.

EmbML · S4 · Hardware & deployment42 / 47

Threat model

Edge deployment improves data privacy and worsens model confidentiality

ThreatWhat the attacker getsMitigation
Model extractionthe weights, by reading flash or by queryingSecure boot, encrypted flash, read-out protection, rate limiting
Membership inferencewhether a record was in the training setDifferential privacy; avoid over-fitting
Adversarial inputscontrolled misclassificationAdversarial training, input plausibility checks
Physical / sensor spoofinginjected signals treated as real — ultrasonic audio, projected lightMulti-sensor consistency, band-limiting in the analog chain
Power / EM side channelarchitecture and sometimes weights, from the current traceConstant-time kernels, masking, noise injection — all costly
Fault injectionskipped instructions, corrupted resultsRedundant computation, output plausibility checks

The side-channel row is the one specific to this domain: physical access means the current trace is correlated with the computation.

EmbML · S4 · Hardware & deployment43 / 47

Frontier

Language models at the edge are a quantization story first

EmbML · S4 · Hardware & deployment44 / 47

Language models at the edge

Be precise about what exists

Real

  • 1–3 B parameter models on phones and embedded modules
  • 4-bit weight quantization — AWQ, GPTQ
  • KV-cache management, speculative decoding
  • Small task-specific models distilled from large ones

Not real

  • Language models on kilobyte-class microcontrollers
  • The binding constraint is weight memory and bandwidth per generated token
  • No amount of clever kernel work removes it
  • The productive embedded pattern is still Session 1’s cascade
EmbML · S4 · Hardware & deployment45 / 47

Session 5

Presentations — what to prepare

One landmark paper per presenter, one topic area each (A quantization · B pruning and distillation · C efficient architectures and NAS · D TinyML systems and hardware). 15 minutes + 5 minutes discussion.

Why: Claim in one sentence with its scope; the mechanism re-derivable; the load-bearing experiment; one real limitation; what it means on a constrained device.

Slides to the lecturer the evening before. Discussants: one prepared question.

EmbML · S4 · Hardware & deployment46 / 47

If you remember one thing

Ask where the bytes come from before asking how fast the arithmetic is. The roofline says which layers an accelerator can help, the toolchain says which it can run at all, and only a measurement on the device, with its method stated, says what you actually have.

EmbML · S4 · Hardware & deployment47 / 47

Before you go

Think about these

  1. An NPU vendor reports 4 TOPS/W. When could the same network be more energy-efficient on the MCU core alone?
  2. Why do depthwise convolutions run at a small fraction of an accelerator's peak? What would a hardware designer change?
  3. Which post-deployment problem — drift, updates, security — is least addressed by research, in your view?
S4 · Hardware & deployment