What the session is for
Reading primary literature critically is the core skill of a PhD. The aim is not to re-teach a paper, but to state precisely what it claims, show the one piece of evidence that carries the claim, explain the mechanism in the cost vocabulary of this course — memory, operations, data movement, energy — and say where its limits lie and what it means for embedded deployment.
How to choose
- One paper per presenter, one topic area per presenter. With four presenters the session covers all four areas A–D; with three, one area is left out.
- First come, first served. Send your first and second choice (e.g. "B3, then D3") to the lecturer by the end of Session 2. Choices are confirmed within a few days.
- Difficulty. Papers are tagged foundational — self-contained, a good choice if machine learning is new to you — or advanced — assumes the background of Session 3 and some comfort with optimisation or hardware. Both are assessed on the same rubric; nobody is penalised for choosing a foundational paper.
- Your own paper? A different landmark paper in the same area may be proposed, if it is peer-reviewed, widely cited, and agreed with the lecturer by the same deadline.
- Discussants. Each presenter is also the discussant for one other talk: read that paper too, and open the discussion with one prepared question.
Topic A — Quantization
Topic A · background: Session 3, Part F
Fewer bits per number
From the integer-only arithmetic that every embedded toolchain uses today, through binary networks and learned quantizers, to the post-training methods that made 4-bit large language models practical.
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Defines the affine (scale and zero-point) quantization scheme, integer-only inference with 32-bit accumulators and fixed-point requantization, and a training procedure that simulates quantization — the scheme behind TensorFlow Lite and most microcontroller runtimes.
FocusDerive the requantization step on one slide. What do the measured latency–accuracy curves on phone CPUs show that a MAC count would not?
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Binarises weights, and then also activations, with real-valued scaling factors, so that convolutions become XNOR and bit-count operations. The extreme end of quantization; read alongside BinaryConnect (Courbariaux, Bengio and David, NIPS 2015, arXiv:1511.00363).
FocusWhere does the accuracy loss on ImageNet come from, and how much of the claimed speed-up would survive on a real microcontroller or NPU?
- Learned Step Size Quantization
Makes the quantizer's step size a trainable parameter, with a carefully derived gradient through the rounding operation and a gradient-scaling rule; reports 2-, 3- and 4-bit ImageNet networks close to, and at 3 bits matching, full-precision accuracy.
FocusExplain the straight-through estimator and the step-size gradient. Why is the gradient scale factor needed, and what happens without it?
- Up or Down? Adaptive Rounding for Post-Training Quantization (AdaRound)
Shows that rounding each weight to the nearest grid point is not optimal for the network's loss, derives a layer-wise objective from a second-order Taylor expansion, and learns the rounding direction from a small unlabelled calibration set — 4-bit weights without fine-tuning.
FocusWalk through the argument from the Taylor expansion to the layer-wise reconstruction loss. Which assumptions does each step make?
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Identifies systematic outlier features that emerge in large transformers and break per-tensor int8 quantization, and handles them with vector-wise scales plus a mixed-precision decomposition that keeps the few outlier dimensions in 16-bit.
FocusConnect the outlier phenomenon to the per-channel versus per-tensor argument of Session 3. What does the decomposition cost in latency?
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
One-shot weight quantization to 3–4 bits using approximate second-order (Hessian) information, made fast enough to quantize models with over a hundred billion parameters in a few GPU-hours.
FocusWhat is the per-layer objective, and how does the algorithm compensate later weights for the error of earlier ones? Compare with AdaRound.
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Observes that a small fraction of weight channels, identified by activation magnitude, dominate the quantization error, and protects them by per-channel scaling rather than mixed precision — hardware-friendly 4-bit weight-only quantization, with an inference system for edge GPUs.
FocusWhy is scaling preferable to keeping salient weights in higher precision on real hardware? What does "weight-only" quantization save, and what not?
Background reading for anyone in this topic: M. Nagel et al., "A White Paper on Neural Network Quantization," 2021, arXiv:2106.08295; R. Krishnamoorthi, "Quantizing deep convolutional networks for efficient inference: A whitepaper," 2018, arXiv:1806.08342.
Topic B — Pruning, sparsity and distillation
Topic B · background: Session 3, Parts G and H
Fewer numbers, and smaller models taught by larger ones
The papers that established pruning as a field, the debate about what pruning actually finds, and the distillation objective that lets a small network inherit a large one's behaviour.
- Learning both Weights and Connections for Efficient Neural Networks
The train–prune–retrain recipe with magnitude pruning, reducing the parameters of AlexNet by 9× and VGG-16 by 13× without accuracy loss on ImageNet. The starting point of modern pruning research.
FocusWhich layers can be pruned most and why? Do 13× fewer parameters mean 13× less memory traffic or time — on what hardware?
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Chains pruning, weight sharing by k-means clustering and Huffman coding to shrink AlexNet's storage by 35× and VGG-16's by 49×, and measures the effect on speed and energy of sparse matrix–vector products on CPU, GPU and mobile GPU.
FocusSeparate the storage gain from the compute gain. Which of the three stages contributes what, and which would survive on a microcontroller?
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Proposes that dense networks contain sparse subnetworks that, reset to their original initialisation, train to full accuracy in isolation; demonstrates it with iterative magnitude pruning and rewinding on small vision benchmarks.
FocusWhat exactly is the hypothesis, and what would falsify it? What changes on larger models (late rewinding), and does any of this make inference cheaper?
- Rethinking the Value of Network Pruning
Finds that, for structured pruning, training the pruned architecture from scratch usually matches fine-tuning the inherited weights — suggesting that pruning is valuable as architecture search rather than as weight selection. The natural counterpart to the lottery ticket paper.
FocusPresent the controlled experiment that carries the claim. How do the two papers' settings differ, and can both be right?
- Pruning Filters for Efficient ConvNets
Removes whole filters (output channels) ranked by their L1 norm, producing a smaller dense network that needs no sparse kernels — the simplest form of structured pruning, with layer-by-layer sensitivity analysis on CIFAR-10.
FocusWhy does structured pruning translate into speed when unstructured pruning does not? Show the sensitivity analysis and what it says about which layers matter.
- Distilling the Knowledge in a Neural Network
Trains a small model to match the temperature-softened output distribution of a large model or ensemble, transferring the "dark knowledge" in the relative probabilities of wrong classes; experiments on MNIST, speech recognition and a large image dataset with specialist models.
FocusDerive why the soft-target loss is scaled by T². What information do soft targets carry that one-hot labels do not, and when is there none?
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Prunes GPT-family models with over a hundred billion parameters to 50–60 % sparsity in one shot, without retraining, by solving a layer-wise sparse reconstruction problem approximately; also handles the 2:4 semi-structured pattern supported by recent GPUs.
FocusWhy does retraining-free pruning become essential at this scale? Relate the 2:4 pattern to the "when does sparsity pay" analysis of Session 3.
Background reading: D. Blalock, J. J. Gonzalez Ortiz, J. Frankle, J. Guttag, "What is the State of Neural Network Pruning?," MLSys 2020, arXiv:2003.03033; A. Romero et al., "FitNets: Hints for Thin Deep Nets," ICLR 2015, arXiv:1412.6550; S. I. Mirzadeh et al., "Improved Knowledge Distillation via Teacher Assistant," AAAI 2020, arXiv:1902.03393.
Topic C — Efficient architectures and neural architecture search
Topic C · background: Session 3, Part D
Designing the network to be cheap from the start
The architecture families that run on phones and microcontrollers, the argument that FLOPs are the wrong target, and the methods that automated the design.
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Builds a whole network from depthwise-separable convolutions and introduces two knobs — a width multiplier and a resolution multiplier — that trade accuracy for cost smoothly. The architecture behind most microcontroller vision models.
FocusDerive the cost ratio of a depthwise-separable versus a standard convolution. How do the two multipliers act on parameters, MACs and activation memory differently?
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
Introduces the inverted residual block — expand, filter depthwise, project back without a non-linearity — and argues for linear bottlenecks from a manifold perspective; evaluated on classification, detection (SSDLite) and segmentation.
FocusWhy invert the classical bottleneck? What does the expanded middle tensor mean for peak SRAM on a microcontroller?
- ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
Shows that FLOPs are an indirect and often misleading proxy for speed, derives four practical guidelines from memory-access cost and parallelism, and measures them on GPU and ARM platforms.
FocusPresent one guideline with its derivation and its measurement. Which of the four would you expect to hold on a microcontroller, and which not?
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Scales depth, width and input resolution together with a single compound coefficient, starting from a baseline found by architecture search; the resulting family dominated the accuracy–parameter trade-off on ImageNet at publication.
FocusWhy should the three dimensions be scaled together? Does the gain in parameters and FLOPs translate into latency and memory on edge hardware?
- Neural Architecture Search with Reinforcement Learning
A recurrent controller generates architecture descriptions and is trained with policy gradients to maximise the validation accuracy of the networks it proposes — the paper that started modern NAS, at a compute cost of hundreds of GPUs.
FocusDecompose the method into search space, search strategy and performance estimation. Which of the three made it so expensive?
- DARTS: Differentiable Architecture Search
Relaxes the discrete choice of operations into a softmax-weighted mixture, so that architecture and weights can be optimised jointly by gradient descent as a bilevel problem — reducing search cost from thousands of GPU-days to a few.
FocusExplain the bilevel optimisation and its first-order approximation. What are the known failure modes (e.g. collapse to skip connections), and why do they occur?
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
Trains a single super-network whose sub-networks of varying depth, width, kernel size and resolution all work, using progressive shrinking; a specialised model for each new hardware target is then selected without retraining.
FocusHow does progressive shrinking prevent sub-networks from interfering? What does the approach assume about the latency predictor for a new device?
Further candidates in this topic: A. Howard et al., "Searching for MobileNetV3," ICCV 2019, arXiv:1905.02244; M. Tan et al., "MnasNet: Platform-Aware Neural Architecture Search for Mobile," CVPR 2019, arXiv:1807.11626; H. Cai, L. Zhu, S. Han, "ProxylessNAS," ICLR 2019, arXiv:1812.00332.
Topic D — TinyML systems and hardware
Topic D · background: Session 4
The machines, the runtimes and the measurements
Landmark accelerators, the systems that made deep learning run on microcontrollers, the benchmark that measures them, and the first steps towards learning on the device.
- In-Datacenter Performance Analysis of a Tensor Processing Unit
The first TPU: a 256×256 systolic array of 8-bit multiply–accumulators with large on-chip buffers, evaluated on Google's production inference workloads against contemporary CPUs and GPUs — and analysed with the roofline model.
FocusUse the paper's rooflines to explain which workloads benefit and why. What carries over from a 40-W datacentre chip to a milliwatt edge NPU?
- Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks
A spatial accelerator built around the row-stationary dataflow, which maximises reuse of weights, inputs and partial sums in local storage to minimise energy-expensive data movement; a fabricated chip running AlexNet's convolutional layers at a few hundred milliwatts.
FocusExplain row-stationary against the other dataflows of Session 4. Where does the energy go, according to the paper's own breakdown?
- MCUNet: Tiny Deep Learning on IoT Devices
Co-designs the network (TinyNAS, which first shrinks the search space to fit the device) and the inference engine (TinyEngine, with memory-aware scheduling and code generation) to run ImageNet-scale classification on an off-the-shelf microcontroller.
FocusWhich gains come from the network and which from the engine? How is peak SRAM, rather than parameter count, made the optimisation target?
- TensorFlow Lite Micro: Embedded Machine Learning for TinyML Systems
The design of the most widely used microcontroller inference framework: an interpreter rather than a code generator, a single statically planned memory arena, and a mechanism for vendor-optimised kernels — and the measured overheads of these choices.
FocusArgue for and against interpretation versus code generation on an MCU. What does the arena planner do, and what does it cost?
- MLPerf Tiny Benchmark
An industry-standard benchmark for ultra-low-power ML: four tasks (keyword spotting, visual wake words, image classification, anomaly detection) with reference models, quality targets, and a methodology for measuring latency and energy on real devices.
FocusWhy these four tasks and models? Explain the energy measurement set-up and the closed versus open divisions. What would you add?
- On-Device Training Under 256KB Memory
Makes fine-tuning possible on a microcontroller with 256 kB of SRAM by combining quantization-aware scaling for stable int8 training, sparse updates of only the most useful layers and tensors, and a training engine that prunes the backward graph at compile time.
FocusWhere does the memory of training go (Session 4's on-device learning figure), and which of the three techniques removes which part?
- Loihi: A Neuromorphic Manycore Processor with On-Chip Learning
Intel's research chip for spiking neural networks: many asynchronous neuromorphic cores with programmable synaptic learning rules on chip, evaluated on sparse-coding problems against a conventional processor.
FocusWhat computational model does the chip implement, and what does it require of the algorithm? Assess the benchmark: is the comparison with a CPU the right one?
Further candidates: L. Lai, N. Suda, V. Chandra, "CMSIS-NN," 2018, arXiv:1801.06601; J. Lin et al., "MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning," NeurIPS 2021, arXiv:2110.15352; T. Chen et al., "TVM," OSDI 2018, arXiv:1802.04799.
How to read a systems paper
Every presenter should read their paper in the four passes below, and check it against the list of common failures that follows. The same method is useful for every paper you will review during your PhD.
How to read a systems paper in four passes method
- Locate the claim (5 min). Title, abstract, conclusion, figure captions. Write the claim as one sentence with an explicit scope: on these models, on this hardware, under this metric. Most misreading happens here, by dropping the scope.
- Locate the evidence (15 min). Go to the tables before the method. What is the baseline, and is it a strong one? What is held constant? Is accuracy measured on the same split as the tuning? Are the compared systems given equal engineering effort — a hand-tuned proposal against an untuned baseline is not a comparison.
- Reconstruct the mechanism (30 min). Now read the method, and try to explain why it should work before accepting that it does. If you cannot, either the paper is unclear or you have found the interesting question.
- Find the boundary (10 min). What would break this? Which item of the checklist below is doing the work? Papers rarely state their own limits fully; the limits are usually visible in what was not tested.
The checklist for compression papers specifically
- Compression of what? Model file size, parameters, MACs, latency or energy — these differ by orders of magnitude and papers slide between them.
- Measured or counted? A FLOP reduction is a count. A speedup is a measurement, and needs a device, a compiler, a clock and a repetition count.
- Which baseline? Compression against an over-parameterised model is easy. The honest comparison is against the best small model trained directly — the "train a small one from scratch" control.
- Accuracy at what operating point? A single accuracy number hides the trade-off curve. Demand the curve.
- How much retraining? A method requiring full retraining is not comparable to one requiring 128 calibration images, whatever their accuracies.
- Does the gain survive composition? Pruning and quantization each recover most accuracy; applied together they may not.
- Variance. One seed is an anecdote. Compression methods are notoriously seed-sensitive at high compression ratios.
Eight ways an efficiency claim goes wrong
- Metric substitution. A FLOP reduction reported as a speedup, or a storage reduction reported as a memory reduction. Ask which of the five budgets (§1.6) actually moved.
- Asymmetric effort. The proposed method is hand-tuned; the baseline is the reference implementation.
- Weak baseline by construction. Compressing an over-parameterised model is easy. The control is "the best small model trained directly".
- Single operating point. One accuracy/size pair hides the curve. Two methods can cross.
- Undisclosed retraining. Methods needing full retraining and methods needing 128 unlabelled images are not comparable at equal accuracy.
- Hardware mismatch. A result measured on a GPU says little about a Cortex-M, and vice versa. Ask what changes with the memory hierarchy.
- Missing variance. One seed, one run, no dispersion. Compression at high ratios is seed-sensitive.
- Non-composability. Two techniques that each recover accuracy may not when combined. Very rarely tested.
What the talk should contain
Fifteen minutes is short. The structure below fits it; slide counts are a guide, not a rule.
- The problem and the claim (1–2 slides). What problem, why it mattered at the time, and the claim in one sentence with its scope.
- The idea (2–3 slides). The mechanism, explained so that the audience could re-derive it. One equation or one figure that carries it.
- The evidence (2–3 slides). The experiment that carries the claim — the load-bearing table or plot — with its baseline and its measurement method.
- The critique (1–2 slides). Apply the checklist above: what is counted versus measured, what baseline, what is not tested, what would falsify the claim.
- The embedded view (1 slide). What it would take to use this on a microcontroller or an edge NPU, in terms of the five budgets of Session 1.
- Since then (1 slide). Two or three follow-up works or how the idea is used in practice today.
Do not re-teach background the course already covered; refer to it. Use the paper's own figures where they help, with attribution, and redraw them when they do not.
Assessment
| Criterion | Weight | Excellent | Inadequate |
|---|---|---|---|
| Claim identified precisely | 20 % | One sentence, correct scope; distinguishes what is shown from what is suggested | Repeats the abstract; scope silently widened |
| Mechanism explained | 20 % | The audience could re-derive the core idea; explained in the course's cost vocabulary | Paraphrases the method section |
| Evidence assessed | 20 % | Names the load-bearing experiment; evaluates baseline strength and measurement method | Accepts headline numbers |
| Critique and embedded relevance | 20 % | Identifies a real limitation and states what it means for deployment on a constrained device | Generic criticism ("more datasets would be good") |
| Delivery and discussion | 20 % | Within time; one figure carries the argument; answers questions precisely; discussant question prepared | Overruns; reads slides; no engagement with questions |
Table 1 — Presentation rubric. Severity of criticism earns nothing on its own; precision and discriminating power do. The weights are a default and may be adjusted by the lecturer.
Timeline
| When | What |
|---|---|
| Session 1 | Topics presented; read this page and shortlist two papers. |
| End of Session 2 | First and second choice sent to the lecturer; papers and discussant pairs confirmed shortly after. |
| Between Sessions 3 and 4 | Optional 10-minute consultation: bring your one-sentence claim and the load-bearing figure. |
| Evening before Session 5 | Slides (PDF) sent to the lecturer; discussants receive them. |
| Session 5 | Talks and discussion, timed as in the figure above. |
Table 2 — Presentation timeline. Exact dates are announced in class.