Embedded Machine LearningPhD course

PhD course · four lectures × 90 minutes + one presentation session

Embedded Machine Learning

How learned models are made to run inside kilobytes, milliwatts and milliseconds: the embedded system and its budgets, the signal and its features, classical and deep models and their compression, and the hardware and toolchains they run on. For doctoral students from machine learning, electrical engineering, computer science, physics and the life sciences.

LevelPhD / doctoral school
Format4 × 90 min lectures + 1 × 90 min presentations
LanguageEnglish
Prerequisitescalculus, linear algebra, some programming
Assessmentpaper presentation

About the course

Most machine-learning teaching optimises one number: accuracy on a held-out set. Embedded machine learning cannot. A model that reaches 94 % accuracy but needs 900 kB of working memory does not "almost work" on a microcontroller with 320 kB — it does not run at all. Here feasibility is a hard set of constraints, and the interesting intellectual content lies in how accuracy trades against memory, latency, energy and engineering effort along that constraint boundary.

The course is written for doctoral students whose research touches sensing, signal processing, wearable or industrial devices, robotics, or efficient machine learning. Participants arrive with very different backgrounds, so the material has two layers. The lecture covers the core in four sessions of 90 minutes. The web chapters go further: they contain primer sections that build the machine-learning and hardware background from first principles for those who need it, and depth sections with derivations and research-level detail for those who want it.

SESSION 1 Embedded systems & the case for edge ML · budgets, constraints · cost of data movement · apps & workloads · the ML workflow SESSION 2 Data, preprocessing & features · sensors, sampling, ADC · filtering, normalising · STFT, mel, MFCC, IMU · select & reduce SESSION 3 Classical vs deep models & compression · ML & NN primer · classical cost models · efficient archs, NAS · quantise, prune, KD SESSION 4 Hardware platforms & deployment · CPU, memory, roofline · MCU · DSP · GPU · NPU · FPGA · analog · neuro · toolchains, MLOps four lectures × 90 min — the order follows the pipeline from sensor to silicon Session 5 — student presentations: one landmark paper each, four topic areas quantization · pruning & distillation · efficient architectures & NAS · TinyML systems & hardware every session is priced against the same five budgets: flash · SRAM · MACs · latency · energy per inference
Fig 1The course in one picture. Four lecture sessions follow the embedded ML pipeline from the device and its budgets, through the signal and its representation, to the model and finally the silicon it runs on. A fifth session is given over to student presentations of landmark papers from the four topic areas listed in the copper band.

Sessions

Session 1 · 90 min

Embedded systems and the case for edge ML

What is scarce on a small device, why it is still worth running models there, and the five budgets every design is priced against.

  • microcontrollers in plain terms
  • an ML primer: model, training, inference
  • applications, workloads, sensor fusion
  • energy of data movement; the cascade

Session 2 · 90 min

Data acquisition, preprocessing and features

From physical quantity to feature vector: what each stage of the chain discards, and the cheapest representation that keeps what the model needs.

  • sensors; time and frequency
  • sampling, aliasing, quantization, noise
  • datasets, normalisation, leakage
  • STFT, mel, MFCC, IMU features; PCA, LDA

Session 3 · 90 min

Classical versus deep learning, and compression

Model families and their inference cost, built up from a primer; efficient architectures; quantization, pruning and distillation.

  • loss, gradient descent, evaluation
  • SVMs, trees, neural networks, CNNs
  • MobileNets, EfficientNet, NAS
  • compression (page 2)

Session 4 · 90 min

Hardware platforms and deployment

What the processor does with a network; the platform landscape from MCU to neuromorphic; and the path from a trained graph to a monitored binary.

  • memory hierarchy, roofline, GEMM
  • MCU, DSP, GPU, NPU, FPGA, analog
  • compilers, memory planning, benchmarks
  • drift, on-device learning, security

Session 5 · 90 min

Paper presentations

Three or four participants each present a landmark paper from a different topic area, followed by discussion.

  • A — quantization
  • B — pruning, sparsity, distillation
  • C — efficient architectures and NAS
  • D — TinyML systems and hardware

Learning outcomes

On completing the course a participant should be able to:

  1. Budget. From a device datasheet and an application specification, derive the admissible model size, peak working memory, operation count and energy per inference, and say which of them binds.
  2. Represent. Specify sampling rate, resolution, windowing and features for a sensing task, and justify each choice by what it keeps and what it costs.
  3. Choose. Select between a classical and a deep model for a stated task, data regime and budget, and evaluate it honestly under class imbalance and subject-wise splits.
  4. Compress. Explain and apply quantization, pruning and distillation, and predict which gains a given hardware target can actually realise.
  5. Analyse. Place a layer on the roofline of a named device and predict whether it is compute- or memory-bound.
  6. Deploy. Describe the path from a trained graph to a running binary, design an honest benchmark, and plan for drift and updates after deployment.
  7. Read. Critically present a primary paper in the field: its claim, its evidence, its mechanism, its limits and its relevance to constrained devices.

How to use these materials

Each session has a web chapter (this site) and a slide deck (PDF and HTML, linked from each chapter and from the cards above). The slides follow the lecture; the chapters contain more than 90 minutes can hold, deliberately. Sections marked depth go further than 90 minutes allow: the lecture presents their result, the page gives the full argument. Sections marked method describe a procedure to follow. Recurring boxes carry different kinds of content:

In one sentenceThe session compressed to its single load-bearing idea. Read it first; read it again afterwards.
FoundationsBackground some participants will not have — CMOS power, decibels, the Fourier transform, two's complement, caches. Self-contained; skip if you know it.
DerivationThe mathematics done properly rather than quoted.
Common misconceptionA specific, widespread wrong belief, and what is actually true.
InteractiveA live calculator or plot. Several results in this course are easier to feel than to read: move the sliders.
TakeawayA closing card per session. Before Session 5, read the four takeaways and the four "in one sentence" boxes.

Primer sections for participants without a machine-learning background: what machine learning adds (Session 1), signals in time and frequency (Session 2), and Part A of Session 3. Reading them before the corresponding lecture is the single most effective preparation.

Assessment

The course is assessed by a paper presentation in Session 5. Each presenter chooses one landmark paper from a different topic area, presents it in 15 minutes and leads a 5-minute discussion, and acts as discussant for one other talk. The presentations page lists the topics and papers, explains how to read a systems paper, and gives the rubric. Choices are due at the end of Session 2.

A note on numbers

Figures quoted from papers are attributed to their source and reflect the hardware and software of that paper's date. Hardware figures quoted without a citation are order-of-magnitude teaching values, not datasheet guarantees: always re-derive against the datasheet of the part you actually have. Where a worked example rests on assumptions, the assumptions are stated next to it.