Embedded Machine Learning · PhD course
Session 03
A primer on learning, the two model families and their inference cost, the architectures designed for small devices, and the three ways to shrink a trained network.
The question of the session
A model is a parameterised function fitted to examples — and, on a device, a bill in bytes and operations.
01
Loss, data splits, gradient descent, evaluation — the minimum to read the rest of the course.
Primer
Primer · learning as optimisation
The first term is the average training loss; the regulariser R (e.g. weight decay ‖θ‖²) prefers simpler parameters, with strength λ chosen by the engineer — a hyper-parameter.
Primer · generalisation
Primer · honest evaluation
Primer · optimisation
Primer · evaluation
The metric mistake that ships products
At 90 % recall this detector raises about nine false alarms for every true one — with an AUC of 0.97 and 92 % accuracy.
02
When a non-neural model still wins — and what each family costs at inference.
The thesis
Choosing a model family is choosing which shapes are cheap to express — and which dimension your inference cost is allowed to grow with.
The regime question
They lose when the input is high-dimensional and the useful features are unknown.
Inductive bias made visible
Logistic regression cannot separate these classes at all — until you add the feature x²+y², which is the entire point of the kernel trick.
The spine of the session
| Model | Arithmetic | Parameters | RAM | Remark |
|---|---|---|---|---|
| Logistic regression | C·D MACs | C(D+1) | O(C) | The reference baseline |
| Gaussian naive Bayes | C·D | 2CD | O(C) | Robust with tiny data |
| k-NN, brute force | N·D MACs | N·D | O(k) | No training; prohibitive memory |
| Linear SVM | D MACs per binary classifier | D+1 | O(1) | One-vs-rest: same cost as logistic, different loss |
| Kernel SVM (RBF) | n_SV·D + n_SV exp | n_SV·D | O(n_SV) | Grows with the dataset |
| Decision tree, depth d | d compares | ~2ᵈ nodes | O(1) | Compiles to nested if |
| Random forest, T trees | T·d compares | T·2ᵈ | O(1) | Thresholds quantize well |
| Boosted stumps, T rounds | T compares + T adds | ~3T | O(1) | Best accuracy-per-byte on tabular features |
Two of these cost comparisons, not multiplications — so a MAC accelerator does nothing for them.
Fitted, not sketched
Delete any non-circled point and the boundary does not move. That sparsity is the beautiful part — and the fatal part.
Why this matters on a microcontroller
A boosted ensemble compiles into straight-line comparison code with the thresholds as immediates.
No matrix, no accumulator, no arena. Tens of comparisons and adds, often under a microsecond, with a completely flat memory profile. For tabular sensor features this is frequently both the most accurate and the cheapest option.
03
Count the cost at every step.
Primer · neurons
Primer · activations
Convolutional networks
Read three costs off one picture
MACs = H·W·k²·C_in·C_out · params = k²·C_in·C_out · activations = H·W·C_out
Learn these
| Layer | Parameters | MACs | Output activations |
|---|---|---|---|
| Dense, D_in → D_out | D_in·D_out + D_out | D_in·D_out | D_out |
| Conv 2-D, k×k, C_in→C_out | k²·C_in·C_out + C_out | H·W·k²·C_in·C_out | H·W·C_out |
| Depthwise conv, k×k | k²·C_in | H·W·k²·C_in | H·W·C_in |
| Pointwise 1×1 | C_in·C_out | H·W·C_in·C_out | H·W·C_out |
| Grouped conv, g groups | k²·C_in·C_out/g | H·W·k²·C_in·C_out/g | H·W·C_out |
| Dilated conv, rate r | same as conv | same as conv | same as conv |
| Batch norm (inference) | 2C stored | 0 after folding | in place |
| Residual add | 0 | H·W·C adds | holds the skip tensor live |
Learn these; the rest of the course assumes them. Note the dilated row: larger receptive field, identical cost.
Derivation
ratio = (k²·C_in + C_in·C_out) / (k²·C_in·C_out) = 1/C_out + 1/k² ≈ 1/9 for a 3×3 kernel.
Folding, and why it is mandatory
Zero inference cost and one fewer pass over the activations. The fold is also a PRECONDITION for quantization — an unfolded BN would need a floating-point per-channel scale at runtime, and the fold is what lets per-channel weight scales absorb it instead.
Decision guide
Break
Five minutes.
Next: efficient architectures, then compression.
04
Designing the network to be cheap from the start — presentation topic C.
The thesis
Every efficient architecture is a factorisation plus an assumption.
The design work is choosing a factorisation whose assumption survives on your hardware. Search is a way of automating that choice against a measured cost model.
Published numbers
Sources: MobileNets Tables 6 & 8; MobileNetV3 Table 3; ShuffleNet V2 Table 8; EfficientNet Table 2.
A memory argument, not an accuracy argument
The final projection is linear — no ReLU — because a ReLU in a low-dimensional space destroys information that cannot be recovered.
Derivation
α · β² · γ² ≈ 2 — because depth enters the cost linearly while width and resolution enter quadratically.
Decomposed
Without the hardware box you get an accurate model. With it you get a deployable one.
05
Fewer bits, fewer numbers, a smaller student — presentation topics A and B.
The toolkit
| Family | What shrinks | Typical gain | Pays off when |
|---|---|---|---|
| Quantization | bits per weight and activation | 4× memory; 2–4× speed on int8 | almost always — the default |
| Pruning | weights, or whole channels | 2–10× weights; 1.5–3× compute | structured: always; unstructured: only with sparse support |
| Distillation | the model — a student replaces the teacher | whatever the student allows | a good teacher and a training budget |
| Low rank / sharing | rank; distinct values | 2–5× on large dense layers | large dense layers; flash-bound |
The whole idea in one sentence
Quantization replaces a per-value exponent with one scale shared by a whole tensor.
That is why it is so cheap in silicon — and exactly why it breaks whenever the values in a tensor are not of comparable size.
Foundations
An int8 multiply costs roughly a twentieth of an fp32 one — before counting the halved memory traffic.
Derivation
The zero-point z is an INTEGER, and that is not a detail: it guarantees that real zero is represented exactly. Padding introduces zeros everywhere, and an inexact zero would inject a bias at every border and every masked position.
Do it with a number
Maximum representation error = s/2 = 2.16 × 10⁻³, uniform across the range.
Derivation · Jacob et al., CVPR 2018
M is empirically always in (0,1). Write M = 2^(−n)·M₀ with M₀ ∈ [0.5,1) stored as a fixed-point int32: multiplying by M is a widening multiply and a rounding right shift. That is what 'integer-only' means — it runs on a core with no FPU at all.
The failure everyone meets once
The fix is granularity, not more bits: one scale per output channel, a vector of requantization multipliers, no extra inner-loop arithmetic.
Measured on that tensor
Clip when levels are scarce, keep the tail when they are plentiful — and choosing costs nothing at inference.
STE and LSQ
The STE is a biased estimator of a gradient that does not exist — and it works, reliably, which is one of the more interesting unexplained facts in the area. LSQ treats the step size as a trained parameter and made 2–4-bit QAT routine rather than heroic.
The tension
The sparsity pattern that costs the least accuracy is the one your processor cannot exploit.
Everything in this session is a negotiation between those two facts.
Taxonomy
Accuracy per removed weight falls left to right. Realised speedup rises left to right. Choose by which you are short of.
The claim, and the counter-claim
Read the critiques alongside it: results are sensitive to the learning-rate schedule and need rewinding to an early checkpoint on larger models.
Derivation
Compressing one way makes the other way harder: combining quantization with unstructured sparsity is far less attractive than combining it with structured sparsity — which is why most microcontroller deployments use structured pruning only.
And size is the easy part
Skipping a multiply saves ≈0.2 pJ. The index decode, the irregular load and the lost vectorisation typically cost more.
The idea
A label tells the student one bit. A teacher’s softened output tells it how the whole class space is arranged.
“This is stop — and it is much more like up and off than like dog.” That relative structure over the wrong classes is what Hinton called dark knowledge, and the one-hot label never contained it.
Mechanism
Zero inference cost. The student ships alone.
Computed on one set of logits
Entropy H rises with T. At T = 1 almost nothing beyond the label; at T = 8 the useful ranking is washing out too.
Putting it together
Before Session 4
B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” CVPR 2018 (§2–3); S. Williams, A. Waterman, D. Patterson, “Roofline,” CACM 52(4), 2009.
Why: Jacob et al. is the arithmetic of every int8 toolchain; the roofline paper is the mental model for Session 4.
Presenters in topics A–C: today's second half is your background. Re-read the matching part of the web chapter.
If you remember one thing
Pick the model by the data and the budget; count its parameters, activations and operations; then compress in an order the hardware can cash in — int8 last, structure before scatter, distillation to recover — and check every claim on the device.
Before you go