Part EAn overview of compression
A trained network is usually far larger and more precise than it needs to be. Compression removes that excess — but each technique removes a different kind of excess, needs a different amount of retraining, and pays off only on hardware that can exploit it.
3.21Four families, one question
| Family | What shrinks | Typical gain | Retraining | Pays off when |
|---|---|---|---|---|
| Quantization | bits per weight and activation | 4× memory (fp32 → int8), 2–4× speed on int8 SIMD/NPU | none (PTQ) to full (QAT) | almost always — the default first step |
| Pruning | number of weights, or whole channels | 2–10× weights (unstructured); 1.5–3× compute (structured) | fine-tuning, often iterative | structured: always; unstructured: only with sparse kernels or hardware |
| Knowledge distillation | the model itself — a smaller student replaces the teacher | whatever the student architecture allows | train the student | you can afford a training run and have a good teacher |
| Low-rank / weight sharing | rank of weight matrices; distinct weight values | 2–5× on large dense layers; codebook storage | fine-tuning | large dense or attention layers; storage-limited flash |
Table 3.6 — The compression families. Gains are typical orders of magnitude reported in the literature for vision and audio CNNs; they vary widely by model and task.
The techniques compose. The classic pipeline of Deep Compression (Han, Mao and Dally, 2016) prunes, then quantizes the survivors by clustering, then entropy-codes the result; a modern embedded flow is more often "choose an efficient architecture, distil it from a large teacher, prune channels if latency is still too high, and quantize to int8 last". The order matters, because each step changes the statistics the next one relies on — a closing section returns to this.
Part FQuantization
In one sentence
Quantization replaces a per-value exponent with one scale shared by a whole tensor — which is why it is so cheap in silicon, and exactly why it breaks whenever the values in a tensor are not of comparable size.
3.22Number formats
| Format | Bits | Range / precision character | Embedded relevance |
|---|---|---|---|
| float32 | 32 | Huge dynamic range, ~7 decimal digits | Training reference; rarely deployed on MCUs |
| float16 | 16 | Narrow exponent; overflows and underflows easily | GPU inference; needs loss scaling in training |
| bfloat16 | 16 | float32's exponent, 8-bit mantissa | Drop-in for training; little benefit on MCUs without hardware |
| int8 (affine) | 8 | 256 uniformly spaced levels over a chosen range | The workhorse. Universally supported by embedded runtimes |
| int4 / int2 | 4 / 2 | Very coarse; usually weights only | Needs QAT or careful PTQ; strong on LLM weights |
| Qm.n fixed point | 16 / 32 | Fixed binary point, no per-tensor scale | Classic DSP style; still used in hand-written kernels |
| Binary / ternary | 1 / ~1.6 | Sign only | Excellent on FPGAs; large accuracy cost on hard tasks |
Table 3.7 — Formats. Note the structural difference: floating point spends bits on an exponent to get dynamic range per value; integer quantization spends a single shared scale to get dynamic range per tensor. Quantization works precisely because activations and weights within a tensor have similar magnitudes — and it fails exactly when they do not.
Foundations · two's complement and Q-format fixed point
Two's complement is how every integer in this chapter is stored. An N-bit word represents −2N−1 … 2N−1−1; negation is “invert all bits and add one”, which is what makes addition and subtraction the same circuit for signed and unsigned values. Two practical consequences appear in real kernels: the range is asymmetric — int8 runs from −128 to +127, so a symmetric quantizer usually restricts itself to ±127 and wastes one level — and addition wraps silently on overflow, which is why quantized kernels accumulate in int32 and use saturating instructions at the boundaries.
Q-format fixed point is the older sibling of affine quantization. Qm.n means an integer interpreted with n fractional bits, i.e. a scale fixed at 2−n — so Q1.15 covers [−1, 1) in steps of 2−15. Multiplying Q1.15 by Q1.15 gives Q2.30, which is why DSP kernels accumulate in 32 bits and shift down at the end.
Affine quantization is exactly this idea with two generalisations: the scale is an arbitrary real number rather than a power of two, and a zero-point lets the representable range be off-centre. The power-of-two restriction is what let 1990s DSPs implement rescaling as a shift; §3.24's M = 2−nM₀ is the modern compromise that recovers the shift while keeping an arbitrary scale.
3.23The affine map depth
Derivation
We want to represent real values x in a range [α, β] using b-bit integers q ∈ [Qmin, Qmax] by an affine relation
Requiring that α maps to Qmin and β to Qmax gives
The zero-point z is an integer, and that is not a detail: it guarantees that the real value 0 is represented exactly. Zero must be exact because padding introduces zeros, and an inexact zero would inject a bias at every image border and every masked position.
Symmetric quantization sets z = 0 and s = max|x| / (2b−1 − 1). It wastes half the range for a strictly positive tensor such as a post-ReLU activation, but it removes the cross-terms in §3.24 and is therefore preferred for weights. The standard combination in deployment toolchains is symmetric per-channel weights, asymmetric per-tensor activations.
Worked example
A weight tensor spans [−0.62, 0.48]; quantize to int8 asymmetric, so Qmin = −128, Qmax = 127.
Maximum representation error is s/2 = 2.16 × 10⁻³, uniform across the range. Compare with float32's relative error: quantization error is absolute, so small values suffer proportionally far more. This is why the distribution of a tensor, not just its extremes, determines quantization damage.
3.24Integer-only inference, derived depth
The core result (after Jacob et al., CVPR 2018)
Let the two operands and the output be quantized with their own parameters: r1 = s1(q1 − z1), r2 = s2(q2 − z2), r3 = s3(q3 − z3). For an inner product of length N:
Everything inside the sum is an integer operation, accumulated in int32. Expanding the product isolates the terms that can be precomputed:
The first term is the ordinary integer GEMM. The third term depends only on the weights and folds into the bias at compile time. The second is a row sum of the activations scaled by the weight zero-point z₂, computed once per row and shared across output channels. The fourth is a constant. With symmetric weights (z₂ = 0) the second term vanishes entirely — which is why asymmetric activations with symmetric weights cost nothing extra, and why asymmetric weights would cost one extra reduction per row.
Requantization. The only non-integer quantity is M = s₁s₂/s₃, empirically always in (0,1). Write it as
and store M₀ as a fixed-point int32. Multiplying by M is then a widening multiply followed by a rounding right shift by n — pure integer arithmetic, no floating-point unit anywhere in the inference path. This is what "integer-only" means, and it is what allows deployment on cores with no FPU at all.
Where the bias goes. Biases are quantized to int32 with scale s₁s₂ — the same scale as the accumulator — so they add directly into the accumulation with no rescaling. Quantizing a bias to int8 would be a serious error: biases have a much larger dynamic range than weights and are added once, so their storage cost is negligible while their precision cost is not.
3.25Granularity: per-tensor, per-channel, per-group
One scale for a whole tensor assumes its values share a magnitude. For ordinary convolutions this is roughly true. For depthwise convolutions it is false: each channel has its own independent filter, and after batch-norm folding the per-channel weight ranges can differ by two orders of magnitude. A single per-tensor scale sized for the largest channel then quantizes the smallest channels to a handful of levels — sometimes to all zeros — and the network collapses. This is the standard explanation for the well-documented failure of naive per-tensor PTQ on MobileNet-class architectures, and per-channel weight scales (one s per output channel) fix it at negligible cost, since the requantization multiplier simply becomes a vector.
3.26Calibration: choosing the clipping range
| Method | Chooses α, β to… | Behaviour |
|---|---|---|
| Min–max | cover the observed extremes | Safe but a single outlier ruins the scale |
| Percentile (e.g. 99.99 %) | cover a quantile of the distribution | Robust; the percentile is a hyperparameter |
| MSE / L2 | minimise E[(x − x̂)²] over candidate ranges | Principled default; cheap grid search per tensor |
| KL divergence (entropy) | minimise divergence between the float and quantized distributions | TensorRT's classic method; effective on activations |
| Moving average min–max | track ranges over batches during training | The standard for QAT |
Table 3.8 — Calibration criteria. A few hundred representative samples suffice; what matters is that they are representative — calibrating a keyword spotter on silence produces activation ranges that clip every real utterance.
3.27Advanced post-training quantization
Cross-layer equalisation. A positively-homogeneous activation such as ReLU satisfies f(αx) = αf(x) for α > 0. So for two consecutive layers you may scale the output channels of the first by a diagonal S and the corresponding input channels of the second by S−1 without changing the function at all. Choosing S to equalise the per-channel ranges makes the network far more amenable to per-tensor quantization — a data-free transformation that recovers most of the gap on MobileNet-class models (Nagel et al., ICCV 2019).
Bias correction. Quantization error has a non-zero mean per channel, E[ΔW·x] = ΔW·E[x] ≠ 0, which shifts every output. Subtracting this expected shift from the bias — using either real data or the batch-norm statistics already stored in the model — removes a systematic error for free.
AdaRound: rounding is a decision, not a function
Rounding to nearest minimises the error of each weight independently. But we do not care about weight error; we care about the error in the layer's output. A second-order expansion of the task loss around the trained weights shows that the relevant quantity couples weights through the Hessian, so the per-weight-optimal rounding is not the jointly-optimal one. AdaRound (Nagel et al., ICML 2020) therefore learns, per weight, whether to round up or down, by optimising a layer-wise reconstruction objective
where h(V) ∈ [0,1] is a differentiable relaxation of the rounding decision and the regulariser pushes it toward 0 or 1 by the end of optimisation. It needs no labels and only a small calibration set — a striking amount of accuracy recovered for a few minutes of per-layer optimisation, and the reason "PTQ" now means something much stronger than min–max rounding.
3.28Quantization-aware training
When PTQ is not enough — sub-8-bit weights, activation-sensitive architectures, tight accuracy requirements — simulate quantization during training.
Fake quantization. Insert x̂ = s·(clamp(round(x/s)+z) − z) into the forward pass while keeping a float32 master copy of the weights. The forward pass now sees the quantization error, so the optimiser can route around it.
The straight-through estimator
round(·) has zero derivative almost everywhere, so the gradient would vanish. The straight-through estimator (Bengio et al., 2013) simply defines the backward pass of the rounding node to be the identity inside the clipping range and zero outside:
This is a biased estimator of a gradient that does not exist — and it works, reliably, which is one of the more interesting unexplained facts in the area. Note the clipping behaviour: weights pushed outside the range receive no gradient and can never come back, so range selection interacts with trainability.
Learned step size quantization (LSQ). Esser et al. (ICLR 2020) treat the step size s as a trained parameter, differentiating through the quantizer with respect to s:
with a gradient scale of 1/√(N·Qmax) to balance the step-size gradient against the weight gradients. LSQ reaches state-of-the-art accuracy at 2–4 bits and made low-bit QAT routine rather than heroic.
Practicalities. Batch norm must be folded during QAT, not after, or training and inference graphs differ. Freeze the batch-norm statistics part-way through. Start from a trained float model — QAT from scratch is slower and no better. Budget 10–20 % of the original training epochs.
3.29Below 8 bits, and mixed precision
Weights tolerate low precision better than activations, because activations must accommodate input-dependent outliers while weights are fixed after training. Hence the standard asymmetry: 4-bit weights with 8-bit activations (W4A8) is often nearly free with a good method, while 4-bit activations usually requires QAT. Binary and ternary networks (BinaryConnect, XNOR-Net) reduce the multiply to a sign flip and are compelling on FPGAs, at an accuracy cost that is small on easy tasks and large on hard ones.
Mixed precision assigns different bit widths to different layers according to sensitivity. Sensitivity can be measured directly (quantize one layer at a time and record the accuracy drop) or estimated from second-order information — layers with large Hessian eigenvalues are more sensitive, the basis of the HAWQ family of methods. The practical recipe: measure per-layer sensitivity, keep the first and last layers at higher precision (they usually are the most sensitive, and they are usually small), and spend the remaining budget where the sensitivity curve is flattest.
Quantization bugs, in order of how often they occur
- Calibration data unrepresentative of deployment — the single most common cause of "quantization broke my model".
- Per-tensor scales on a depthwise-separable architecture (§3.25).
- Batch norm not folded, or folded after calibration rather than before.
- Biases quantized to 8 bits instead of 32.
- Concatenation of tensors with different scales without re-scaling — a silent accuracy loss with no error message.
- Accuracy measured on the float model but latency on the quantized one, or vice versa.
- Residual additions where the two branches carry different scales.
Diagnose by computing per-layer signal-to-quantization-noise ratio, 10 log₁₀(‖y‖²/‖y − ŷ‖²), on a calibration batch and finding the first layer where it collapses. Debugging quantization is a localisation problem, and this is the localiser.
Part GPruning, sparsity and low rank
In one sentence
Removing weights is easy and removing them usefully is hard: the sparsity pattern that costs the least accuracy is exactly the pattern that no ordinary processor can exploit.
3.30Taxonomy
3.31Saliency: which weights can go depth
Magnitude. Remove the smallest |w|. Justified if all weights have comparable influence, which after batch-norm folding they roughly do within a layer. It is embarrassingly effective and remains the baseline everything is compared against.
Optimal Brain Damage, derived
Expand the loss around a trained (hence near-stationary) parameter vector:
At a minimum the gradient g vanishes, so the first term drops. LeCun et al. (1990) additionally assume the Hessian is diagonal and neglect higher orders. Deleting weight i means δwi = −wi, so the predicted loss increase — the saliency — is
This says something magnitude pruning misses: a large weight in a flat direction is safe to remove, and a small weight in a sharply curved direction is not.
Optimal Brain Surgeon (Hassibi & Stork, 1993) drops the diagonal assumption and additionally allows the remaining weights to be updated to compensate. Solving the constrained problem gives the saliency and the compensating update
which is more accurate but requires the inverse Hessian — infeasible for large models, though modern approximations of exactly this idea (layer-wise, with efficient Hessian estimates) underpin current one-shot pruning and quantization methods for large models.
First-order criteria. When the model is not at a stationary point — during training, or after fine-tuning on new data — the gradient term does not vanish and |giwi| is an effective and cheap saliency, widely used for structured channel pruning.
3.32Schedules, and the lottery ticket
One-shot pruning removes the target fraction at once and fine-tunes. Iterative magnitude pruning alternates prune-and-retrain in small steps and consistently reaches higher sparsity at the same accuracy, because the network is given the chance to redistribute function into the surviving weights.
The standard automated schedule (Zhu & Gupta, 2017) increases sparsity from si to sf with a cubic profile — fast at first, when there is slack, then slow:
The lottery ticket hypothesis (Frankle & Carbin, ICLR 2019) states that a randomly initialised dense network contains a sparse subnetwork which, trained in isolation from the same initialisation, matches the full network's accuracy in comparable time. The procedure is: train, prune by magnitude, reset the survivors to their original initial values, retrain, repeat. On MNIST and CIFAR-10 the winning tickets found were typically 10–20 % of the original size.
The claim matters because it separates two things usually conflated: whether a sparse architecture is sufficient, and whether a particular initialisation of it is trainable. Read the critiques alongside it: results are sensitive to learning-rate schedule and require rewinding to an early-training checkpoint rather than to initialisation on larger models; and Liu et al. ("Rethinking the Value of Network Pruning," ICLR 2019) argue that for structured pruning, retraining the discovered architecture from scratch often matches the inherited weights — which would make pruning a form of architecture search rather than a form of weight selection. This tension is one of the central debates of the pruning literature, and a good subject for a presentation in topic B.
3.33The systems reality: when does sparsity pay?
Break-even sparsity for storage
Store a sparse matrix in compressed-sparse-row form: for each non-zero, one value and one column index; plus one row pointer per row. With bv bits per value and bi bits per index, and a fraction σ of weights surviving, the sparse representation is smaller than the dense one when
For float32 values with 16-bit indices: σ < 32/48 = 0.67, i.e. any sparsity above 33 % saves storage. For int8 values with 8-bit relative indices — the embedded case — σ < 8/16 = 0.5: you need to remove more than half the weights before the sparse format is even smaller, let alone faster. Combining quantization with unstructured sparsity is therefore much less attractive than combining it with structured sparsity, and this arithmetic is the reason most microcontroller deployments use structured pruning only.
And speed is harder than size. Skipping a multiply saves ≈0.2 pJ (Table 1.3); the index decode, the irregular load, and the lost vectorisation typically cost more. General-purpose cores need very high sparsity — commonly cited thresholds are around 80–90 % — before a sparse kernel beats a dense one. This is exactly why N:M semi-structured sparsity was introduced: constraining exactly N of every M consecutive weights to be non-zero gives a fixed, tiny metadata format that hardware can decode at full rate, delivering a real speedup at a modest 50 % sparsity. It is a design in which the statistics were bent to fit the hardware, and it is worth recognising as such.
3.34Structured pruning
Removing whole output channels, filters, attention heads or layers changes the shape of the tensors, so the result is an ordinary smaller network that any runtime executes faster. Criteria in use:
- Filter norm — remove filters with the smallest L1 norm (Li et al., ICLR 2017).
- Batch-norm scale — add an L1 penalty on the BN γ parameters during training; channels whose γ goes to zero contribute nothing and are removed (network slimming). Elegant, because the pruning criterion is trained rather than imposed.
- Taylor / gradient-based — estimate each channel's contribution to the loss directly.
- Learned or searched ratios — AMC (He et al., ECCV 2018) uses reinforcement learning to choose a per-layer sparsity ratio against a hardware latency constraint, which is really a form of architecture search and a natural bridge to neural architecture search (§3.20).
A practical caution: uniform per-layer ratios are almost always wrong. Early layers have few parameters but dominate the MAC count; late layers have most of the parameters and little compute. Pruning uniformly optimises neither budget.
3.35Low-rank and tensor factorisation
When does a low-rank factorisation help?
Replace a dense layer W ∈ ℝDout×Din by W ≈ UV with U ∈ ℝDout×r, V ∈ ℝr×Din. Cost falls from DinDout to r(Din + Dout), so the factorisation is worthwhile only when
For a 1024×1024 layer that means r < 512: you must discard at least half the spectrum before you gain anything. Truncated SVD gives the optimal rank-r approximation of W in Frobenius norm — but the optimal approximation of the weights is not the optimal approximation of the function, so fine-tuning after factorisation is not optional.
For convolution kernels, which are 4-D tensors, the analogous tools are Tucker and CP decompositions; a Tucker-2 decomposition of a k×k×Cin×Cout kernel yields a 1×1 convolution, a smaller k×k convolution and another 1×1 convolution — structurally the same idea as the bottleneck blocks that efficient architectures use by design (§3.17). Note the connection: a well-designed efficient architecture is a factorisation you did not have to discover after the fact.
3.36Weight sharing and coding
Cluster the weights of a layer with k-means into 2b centroids, store a b-bit index per weight plus a small codebook, and fine-tune the centroids by summing the gradients of all weights assigned to each. This is weight sharing, and with entropy coding of the indices it forms the second and third stages of Deep Compression (Han et al., ICLR 2016), which reported 35–49× model-size reduction on AlexNet and VGG-16 without accuracy loss.
The distinction that Deep Compression makes it easy to blur
Huffman coding and weight sharing compress the stored model. They do not make inference faster, because the weights must be decoded before use — and on a microcontroller the decode buffer may not fit. Distinguish carefully:
- Storage compression — smaller firmware image, cheaper over-the-air update. Huffman, weight sharing, codebooks.
- Runtime compression — fewer bytes moved and fewer operations executed at inference. Quantization, structured pruning, factorisation.
Papers that report only the first and readers who assume the second are a recurring source of confusion in this field.
Part HKnowledge distillation
In one sentence
A label tells the student one bit; a teacher’s softened output tells it how the whole class space is arranged — and that extra structure costs nothing at inference, because the teacher is thrown away after training.
3.37The idea
A one-hot label says "this is a keyword". A trained teacher's output distribution says "this is 'stop', it is somewhat like 'up', and nothing at all like 'off'". That relative structure over the wrong classes — Hinton's dark knowledge — encodes the teacher's learned similarity metric, and it is a far denser training signal than a single bit per class. The consequence is that a small student trained on teacher outputs can generalise better than the same student trained on the labels alone, sometimes markedly so, and it costs exactly nothing at inference because the teacher is discarded after training.
3.38The objective, derived depth
Temperature and the T² factor
With teacher logits v and student logits z, define softened distributions pi = softmax(v/T)i and qi = softmax(z/T)i. The distillation loss is
Differentiating the cross-entropy of the softened student against the softened teacher gives
so the soft-target gradient shrinks as 1/T — and after the softening also compresses the logit differences, the effective magnitude falls as 1/T². Multiplying the soft term by T² restores it to the same scale as the hard-label term, so that α means what you think it means and you do not have to retune the learning rate every time you change T.
The high-temperature limit. For large T, expand exp(zi/T) ≈ 1 + zi/T. Assuming logits are zero-meaned per sample, the gradient becomes
i.e. distillation degenerates to least-squares matching of the logits themselves. At low temperature, by contrast, the loss concentrates on the classes the teacher considers plausible and ignores the very negative logits — which is often preferable, since those logits are noisy. The temperature is therefore a knob selecting how much of the teacher's low-confidence structure the student is asked to reproduce.
3.39Variants
| Family | What is matched | Use when |
|---|---|---|
| Response-based (classical KD) | Output distribution | Default; architecture-agnostic |
| Feature-based (FitNets and successors) | Intermediate activations, via a learned projection | Student much deeper/narrower; needs layer pairing |
| Attention transfer | Spatial attention maps derived from activations | Vision; cheap and robust to width mismatch |
| Relational KD | Pairwise distances/angles between samples in embedding space | Metric learning, retrieval, speaker ID |
| Self-distillation / born-again | Same architecture, previous generation as teacher | No larger model available; still often helps |
| Online / mutual learning | Two students teach each other during training | No pre-trained teacher; one training run |
| Data-free KD | Synthesised inputs matching the teacher's BN statistics | Training data cannot be shared — a real constraint in medical and industrial work |
Table 3.9 — Distillation variants. Feature-based methods are more powerful and much fussier: they require choosing which layers to pair and a projection to reconcile dimensions, and a bad pairing actively hurts.
3.40The capacity gap, and the recipe
A stronger teacher is not always a better teacher. Beyond some gap in capacity the student cannot represent the teacher's function, and forcing it to try produces worse results than distilling from a moderate teacher. The standard remedies are a teacher assistant — distil large → medium → small — or simply selecting the teacher by validated student accuracy rather than by teacher accuracy. Report the pairing, not just the student.
Practical recipe
- Train or obtain the teacher. For audio, a pre-trained large model (e.g. a PANNs-class network) is usually better than one you train yourself.
- Start with T ∈ [2,6] and α ∈ [0.5, 0.9]. Tune T first; it matters more.
- Use the same augmentation for teacher and student inputs, so the targets correspond to what the student sees.
- Exploit unlabelled data — it is free supervision under the teacher.
- Validate the student, not the agreement with the teacher. High teacher-agreement with low accuracy means you have distilled the teacher's errors.
Order of operations when combining. The reliable pipeline is: train the teacher → distil into the small architecture → structurally prune with fine-tuning under the distillation loss → quantize (PTQ, or QAT with the teacher still supervising). Distillation is particularly effective as the loss during QAT, because the float teacher provides a stable target while the student's forward pass is being perturbed by fake quantization. Doing it in the other order — quantize then distil — throws away the quantized model's calibration and usually has to be redone.
3.41Putting it together: a compression recipe
Faced with a model that does not meet its budget, the following order of operations is a sound default. It is not the only one, and several presentation papers argue for different orders; treat it as the baseline your own method should beat.
- Fix the budget and the measurement. Flash, peak SRAM, latency and energy on the target device, measured with the deployment toolchain (Session 4). Compression decisions made on proxy counts — parameters, FLOPs — are routinely wrong.
- Choose the architecture first. An efficient family scaled to the budget (Part D) usually beats compressing an inefficient one. If the budget is far away, change the architecture, not the compression ratio.
- Distil while you shrink. Train (or fine-tune) the small model with the original float model as teacher. Distillation works best between floating-point models, before any quantization noise is introduced.
- Remove structure, not scattered weights. If latency or SRAM still binds, prune channels (structured pruning) and fine-tune, keeping the distillation loss; unstructured sparsity only if the target has sparse kernels or hardware support.
- Quantize to int8 last. Post-training quantization with per-channel weights and a calibration set drawn from field-like data. If the accuracy drop exceeds the tolerance, switch to quantization-aware training initialised from the pruned float model — the teacher can stay in the loss.
- Go below 8 bits only with evidence. Mixed precision or 4-bit weights if flash binds and the hardware or kernel library supports them; measure, because sub-byte unpacking can cost more time than it saves.
- Re-measure on the device after every step, and keep the Pareto front of (accuracy, flash, SRAM, latency, energy) rather than a single "best" model.
If you remember one thing from Session 3
Pick the model by the data and the budget; count its parameters, activations and operations; then compress in an order the hardware can cash in: int8 first, structured removal second, distillation to recover, and every claim checked on the device. A compression ratio the target cannot exploit is a number in a paper, not a saving.
Before Session 4
- B. Jacob et al., "Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference," CVPR 2018. arXiv:1712.05877.The paper behind every int8 deployment toolchain; read §2–3 for the arithmetic derived above.
- S. Williams, A. Waterman, D. Patterson, "Roofline: An Insightful Visual Performance Model for Multicore Architectures," Communications of the ACM 52(4), 2009.Short, and the single most useful mental model for Session 4.
Discussion questions
- Quantization replaces a per-value exponent with a per-tensor scale. Give a layer type for which this is benign and one for which it is harmful, and explain why in terms of the value distributions.
- Is pruning a form of weight selection or of architecture search? State the evidence on each side.
- Why does a teacher's "dark knowledge" help a student more on some tasks than on others? Construct a task where it cannot help at all.
- A paper reports 10× compression with no accuracy loss. List the five questions you ask before believing it is 10× cheaper to run.
Exercises
- A tensor has values in [−0.8, 2.4]. Compute scale and zero-point for asymmetric uint8 quantization and for symmetric int8. Quantize and dequantize the value 1.0 in each case and report the error.
- Derive the requantization multiplier for an int8 convolution with input scale 0.02, weight scale 0.005 and output scale 0.05, and express it as a 32-bit fixed-point multiplier and shift.
- A dense layer of 1024×1024 weights is pruned to 90 % unstructured sparsity and stored in CSR format with 16-bit indices. Compare the storage with the dense int8 version. At what sparsity do the two break even?
- A 512×512 weight matrix is replaced by a rank-64 factorisation. Compute the reduction in parameters and MACs, and state when the factorised version could be slower.
Further reading
- M. Nagel et al., "A White Paper on Neural Network Quantization," 2021. arXiv:2106.08295.The best single reference on Part F; read in full at some point.
- R. Krishnamoorthi, "Quantizing deep convolutional networks for efficient inference: A whitepaper," 2018. arXiv:1806.08342.The practitioner's companion, with measured accuracy tables.
- D. Blalock et al., "What is the State of Neural Network Pruning?," MLSys 2020. arXiv:2003.03033.A meta-analysis that every pruning paper should be read against.
- G. Hinton, O. Vinyals, J. Dean, "Distilling the Knowledge in a Neural Network," 2015. arXiv:1503.02531.Nine pages; the origin of the temperature-scaled objective.