EmbML · S3 · Models & compression1 / 55

Embedded Machine Learning · PhD course

Session 03

Classical versus deep learning — and how to make models small

A primer on learning, the two model families and their inference cost, the architectures designed for small devices, and the three ways to shrink a trained network.

90 minutesSession 3 of 5
EmbML · S3 · Models & compression2 / 55

The question of the session

A model is a parameterised function fitted to examples — and, on a device, a bill in bytes and operations.

EmbML · S3 · Models & compression3 / 55

01

A machine-learning primer

Loss, data splits, gradient descent, evaluation — the minimum to read the rest of the course.

EmbML · S3 · Models & compression4 / 55

Primer

The families of machine learning

EmbML · S3 · Models & compression5 / 55

Primer · learning as optimisation

Training minimises a loss over examples

θ* = argmin_θ (1/N) Σᵢ ℓ( f(xᵢ; θ), yᵢ ) + λ·R(θ)
regression ℓ = (ŷ − y)²
classification pᶜ = e^{zᶜ} / Σₖ e^{zₖ} ℓ = −log p_y (cross-entropy)

The first term is the average training loss; the regulariser R (e.g. weight decay ‖θ‖²) prefers simpler parameters, with strength λ chosen by the engineer — a hyper-parameter.

EmbML · S3 · Models & compression6 / 55

Primer · generalisation

Training error falls; test error does not

EmbML · S3 · Models & compression7 / 55

Primer · honest evaluation

Fit on train, choose on validation, report test once

EmbML · S3 · Models & compression8 / 55

Primer · optimisation

Gradient descent, and why feature scaling matters

EmbML · S3 · Models & compression9 / 55

Primer · evaluation

Accuracy hides rare events

EmbML · S3 · Models & compression10 / 55

The metric mistake that ships products

Same detector. One panel lies.

At 90 % recall this detector raises about nine false alarms for every true one — with an AUC of 0.97 and 92 % accuracy.

EmbML · S3 · Models & compression11 / 55

02

Classical models under constraints

When a non-neural model still wins — and what each family costs at inference.

EmbML · S3 · Models & compression12 / 55

The thesis

Choosing a model family is choosing which shapes are cheap to express — and which dimension your inference cost is allowed to grow with.

EmbML · S3 · Models & compression13 / 55

The regime question

Classical wins when at least one of these holds

  • Little data. A few hundred labelled windows: a regularised linear model or a small forest generalises better than a network you cannot regularise into submission.
  • Extreme memory limits. Thirty depth-6 boosted trees fit in a few kilobytes of flash and need almost no RAM.
  • Deterministic latency. A tree’s latency varies with the path; a linear model’s does not vary at all.
  • Interpretability or certification. Medical and automotive regimes are far happier with a model you can print as a rule.
  • Good features already exist. If domain physics gives you the right twenty features, the representation-learning advantage evaporates.

They lose when the input is high-dimensional and the useful features are unknown.

EmbML · S3 · Models & compression14 / 55

Inductive bias made visible

Same data, same features, four different biases

Logistic regression cannot separate these classes at all — until you add the feature x²+y², which is the entire point of the kernel trick.

EmbML · S3 · Models & compression15 / 55

The spine of the session

Inference cost, by family

ModelArithmeticParametersRAMRemark
Logistic regressionC·D MACsC(D+1)O(C)The reference baseline
Gaussian naive BayesC·D2CDO(C)Robust with tiny data
k-NN, brute forceN·D MACsN·DO(k)No training; prohibitive memory
Linear SVMD MACs per binary classifierD+1O(1)One-vs-rest: same cost as logistic, different loss
Kernel SVM (RBF)n_SV·D + n_SV expn_SV·DO(n_SV)Grows with the dataset
Decision tree, depth dd compares~2ᵈ nodesO(1)Compiles to nested if
Random forest, T treesT·d comparesT·2ᵈO(1)Thresholds quantize well
Boosted stumps, T roundsT compares + T adds~3TO(1)Best accuracy-per-byte on tabular features

Two of these cost comparisons, not multiplications — so a MAC accelerator does nothing for them.

EmbML · S3 · Models & compression16 / 55

Fitted, not sketched

Three of sixty-eight points carry the whole decision function

Delete any non-circled point and the boundary does not move. That sparsity is the beautiful part — and the fatal part.

EmbML · S3 · Models & compression17 / 55

Why this matters on a microcontroller

A boosted ensemble compiles into straight-line comparison code with the thresholds as immediates.

No matrix, no accumulator, no arena. Tens of comparisons and adds, often under a microsecond, with a completely flat memory profile. For tabular sensor features this is frequently both the most accurate and the cheapest option.

EmbML · S3 · Models & compression18 / 55

03

Neural networks, from one neuron

Count the cost at every step.

EmbML · S3 · Models & compression19 / 55

Primer · neurons

A neuron, a layer, a network

EmbML · S3 · Models & compression20 / 55

Primer · activations

The non-linearity, and what it costs on integer hardware

EmbML · S3 · Models & compression21 / 55

Convolutional networks

Space shrinks, channels grow — activations front, parameters back

EmbML · S3 · Models & compression22 / 55

Read three costs off one picture

What a convolution actually costs

MACs = H·W·k²·C_in·C_out  ·  params = k²·C_in·C_out  ·  activations = H·W·C_out

EmbML · S3 · Models & compression23 / 55

Learn these

Layer cost formulas

LayerParametersMACsOutput activations
Dense, D_in → D_outD_in·D_out + D_outD_in·D_outD_out
Conv 2-D, k×k, C_in→C_outk²·C_in·C_out + C_outH·W·k²·C_in·C_outH·W·C_out
Depthwise conv, k×kk²·C_inH·W·k²·C_inH·W·C_in
Pointwise 1×1C_in·C_outH·W·C_in·C_outH·W·C_out
Grouped conv, g groupsk²·C_in·C_out/gH·W·k²·C_in·C_out/gH·W·C_out
Dilated conv, rate rsame as convsame as convsame as conv
Batch norm (inference)2C stored0 after foldingin place
Residual add0H·W·C addsholds the skip tensor live

Learn these; the rest of the course assumes them. Note the dilated row: larger receptive field, identical cost.

EmbML · S3 · Models & compression24 / 55

Derivation

One operation split into two cheap ones

ratio = (k²·C_in + C_in·C_out) / (k²·C_in·C_out) = 1/C_out + 1/k² ≈ 1/9 for a 3×3 kernel.

EmbML · S3 · Models & compression25 / 55

Folding, and why it is mandatory

Batch norm disappears into the convolution

y = γ (x − μ)/√(σ²+ε) + β with x = W ∗ a + b
⇒ y = W' ∗ a + b'
W' = γW / √(σ²+ε)
b' = γ(b − μ)/√(σ²+ε) + β

Zero inference cost and one fewer pass over the activations. The fold is also a PRECONDITION for quantization — an unfolded BN would need a floating-point per-channel scale at runtime, and the fold is what lets per-channel weight scales absorb it instead.

EmbML · S3 · Models & compression26 / 55

Decision guide

Classical or deep?

EmbML · S3 · Models & compression27 / 55

Break

Five minutes.

Next: efficient architectures, then compression.

EmbML · S3 · Models & compression28 / 55

04

Efficient architectures

Designing the network to be cheap from the start — presentation topic C.

EmbML · S3 · Models & compression29 / 55

The thesis

Every efficient architecture is a factorisation plus an assumption.

The design work is choosing a factorisation whose assumption survives on your hardware. Search is a way of automating that choice against a measured cost model.

EmbML · S3 · Models & compression30 / 55

Published numbers

Five years moved the frontier an order of magnitude

Sources: MobileNets Tables 6 & 8; MobileNetV3 Table 3; ShuffleNet V2 Table 8; EfficientNet Table 2.

EmbML · S3 · Models & compression31 / 55

A memory argument, not an accuracy argument

Why the residual was inverted

The final projection is linear — no ReLU — because a ReLU in a low-dimensional space destroys information that cannot be recovered.

EmbML · S3 · Models & compression32 / 55

Derivation

Where EfficientNet’s constraint comes from

α · β² · γ² ≈ 2 — because depth enters the cost linearly while width and resolution enter quadratically.

EmbML · S3 · Models & compression33 / 55

Decomposed

Search space, strategy, estimation — and hardware cost

Without the hardware box you get an accurate model. With it you get a deployable one.

EmbML · S3 · Models & compression34 / 55

05

Compression

Fewer bits, fewer numbers, a smaller student — presentation topics A and B.

EmbML · S3 · Models & compression35 / 55

The toolkit

Four families, one question: can the hardware cash it in?

FamilyWhat shrinksTypical gainPays off when
Quantizationbits per weight and activation4× memory; 2–4× speed on int8almost always — the default
Pruningweights, or whole channels2–10× weights; 1.5–3× computestructured: always; unstructured: only with sparse support
Distillationthe model — a student replaces the teacherwhatever the student allowsa good teacher and a training budget
Low rank / sharingrank; distinct values2–5× on large dense layerslarge dense layers; flash-bound
EmbML · S3 · Models & compression36 / 55

The whole idea in one sentence

Quantization replaces a per-value exponent with one scale shared by a whole tensor.

That is why it is so cheap in silicon — and exactly why it breaks whenever the values in a tensor are not of comparable size.

EmbML · S3 · Models & compression37 / 55

Foundations

Two ways to spend bits

An int8 multiply costs roughly a twentieth of an fp32 one — before counting the halved memory traffic.

EmbML · S3 · Models & compression38 / 55

Derivation

Scale and zero-point from a range

x ≈ s · (q − z) de-quantization
q = clamp( round(x/s) + z , Q_min , Q_max ) quantization
s = (β − α) / (Q_max − Q_min)
z = round( Q_min − α / s )

The zero-point z is an INTEGER, and that is not a detail: it guarantees that real zero is represented exactly. Padding introduces zeros everywhere, and an inexact zero would inject a bias at every border and every masked position.

EmbML · S3 · Models & compression39 / 55

Do it with a number

Quantize a weight tensor spanning [−0.62, 0.48]

  1. s = (0.48 − (−0.62)) / 255 = 1.10 / 255 = 4.3137 × 10⁻³
  2. z = round(−128 − (−0.62)/s) = round(−128 + 143.73) = 16
  3. x = 0.31 → q = round(71.86) + 16 = 88, x̂ = s(88 − 16) = 0.3106 — error 5.9 × 10⁻⁴
  4. x = −0.62 → q = round(−143.73) + 16 = −128, clamped exactly
  5. x = 0 → q = 16, x̂ = 0 exactly

Maximum representation error = s/2 = 2.16 × 10⁻³, uniform across the range.

EmbML · S3 · Models & compression40 / 55

Derivation · Jacob et al., CVPR 2018

An inner product with no floating point anywhere

s₃(q₃ − z₃) = Σ_j s₁(q₁ⱼ − z₁) · s₂(q₂ⱼ − z₂)
q₃ = z₃ + M · Σ_j (q₁ⱼ − z₁)(q₂ⱼ − z₂), M = s₁s₂/s₃
Σ_j (q₁ⱼ−z₁)(q₂ⱼ−z₂) = Σ q₁q₂ − z₂ Σ q₁ − z₁ Σ q₂ + N z₁ z₂
↑GEMM ↑row sum ↑folds into bias ↑constant

M is empirically always in (0,1). Write M = 2^(−n)·M₀ with M₀ ∈ [0.5,1) stored as a fixed-point int32: multiplying by M is a widening multiply and a rounding right shift. That is what 'integer-only' means — it runs on a core with no FPU at all.

EmbML · S3 · Models & compression41 / 55

The failure everyone meets once

Why per-tensor scales destroy depthwise layers

The fix is granularity, not more bits: one scale per output channel, a vector of requantization multipliers, no extra inner-loop arithmetic.

EmbML · S3 · Models & compression42 / 55

Measured on that tensor

The right range depends on the bit width

Clip when levels are scarce, keep the tail when they are plentiful — and choosing costs nothing at inference.

EmbML · S3 · Models & compression43 / 55

STE and LSQ

Getting a gradient through a step function

x̂ = s·(clamp(round(x/s)+z) − z) fake quantization in the forward pass
∂x̂/∂x := 1 inside the clipping range, 0 outside (straight-through estimator)
∂x̂/∂s = −x/s + round(x/s) inside, Q_min or Q_max outside (LSQ)

The STE is a biased estimator of a gradient that does not exist — and it works, reliably, which is one of the more interesting unexplained facts in the area. LSQ treats the step size as a trained parameter and made 2–4-bit QAT routine rather than heroic.

EmbML · S3 · Models & compression44 / 55

The tension

The sparsity pattern that costs the least accuracy is the one your processor cannot exploit.

Everything in this session is a negotiation between those two facts.

EmbML · S3 · Models & compression45 / 55

Taxonomy

Three regimes, and what each buys

Accuracy per removed weight falls left to right. Realised speedup rises left to right. Choose by which you are short of.

EmbML · S3 · Models & compression46 / 55

The claim, and the counter-claim

The lottery ticket hypothesis

  • Claim. A randomly initialised dense network contains a sparse subnetwork which, trained from the same initialisation, matches the full network in comparable time.
  • Procedure. Train → prune by magnitude → reset survivors to their original initial values → retrain → repeat.
  • On MNIST and CIFAR-10 the winning tickets found were typically 10–20 % of the original size.
  • Why it matters: it separates two things usually conflated — whether a sparse architecture suffices, and whether a particular initialisation of it is trainable.

Read the critiques alongside it: results are sensitive to the learning-rate schedule and need rewinding to an early checkpoint on larger models.

EmbML · S3 · Models & compression47 / 55

Derivation

Break-even sparsity for storage

CSR: one value + one index per non-zero. σ = fraction SURVIVING
σ·(b_v + b_i) < b_v ⇒ σ < b_v / (b_v + b_i)
float32 + 16-bit index: σ < 32/48 = 0.67 → above 33 % saves
int8 + 8-bit index: σ < 8/16 = 0.50 → remove MORE THAN HALF

Compressing one way makes the other way harder: combining quantization with unstructured sparsity is far less attractive than combining it with structured sparsity — which is why most microcontroller deployments use structured pruning only.

EmbML · S3 · Models & compression48 / 55

And size is the easy part

Speed is harder than size

Skipping a multiply saves ≈0.2 pJ. The index decode, the irregular load and the lost vectorisation typically cost more.

EmbML · S3 · Models & compression49 / 55

The idea

A label tells the student one bit. A teacher’s softened output tells it how the whole class space is arranged.

“This is stop — and it is much more like up and off than like dog.” That relative structure over the wrong classes is what Hinton called dark knowledge, and the one-hot label never contained it.

EmbML · S3 · Models & compression50 / 55

Mechanism

The teacher is discarded after training

Zero inference cost. The student ships alone.

EmbML · S3 · Models & compression51 / 55

Computed on one set of logits

Where the dark knowledge lives

Entropy H rises with T. At T = 1 almost nothing beyond the label; at T = 8 the useful ranking is washing out too.

EmbML · S3 · Models & compression52 / 55

Putting it together

A compression recipe — the baseline your method should beat

  • Fix the budget and the measurement on the target, with the deployment toolchain.
  • Choose an efficient architecture scaled to the budget.
  • Distil from a float teacher while the model is still in floating point.
  • Prune channels if latency or SRAM still binds; fine-tune with the teacher.
  • Quantize to int8 last — PTQ, then QAT if the drop is too large.
  • Re-measure after every step; keep the Pareto front.
EmbML · S3 · Models & compression53 / 55

Before Session 4

Reading

B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” CVPR 2018 (§2–3); S. Williams, A. Waterman, D. Patterson, “Roofline,” CACM 52(4), 2009.

Why: Jacob et al. is the arithmetic of every int8 toolchain; the roofline paper is the mental model for Session 4.

Presenters in topics A–C: today's second half is your background. Re-read the matching part of the web chapter.

EmbML · S3 · Models & compression54 / 55

If you remember one thing

Pick the model by the data and the budget; count its parameters, activations and operations; then compress in an order the hardware can cash in — int8 last, structure before scatter, distillation to recover — and check every claim on the device.

EmbML · S3 · Models & compression55 / 55

Before you go

Think about these

  1. Quantization replaces a per-value exponent with a per-tensor scale. Give a layer for which this is benign and one for which it is harmful.
  2. Is pruning weight selection or architecture search? State the evidence on each side.
  3. A paper reports 10× compression with no accuracy loss. List five questions before believing it is 10× cheaper to run.
S3 · Models & compression