Embedded Machine Learning · PhD course
Session 02
What has already been lost before the model sees the data — and the cheapest representation that keeps what it needs.
The principle
The model can only ever recover what the chain preserved.
Two steps in this session are irreversible — aliasing and clipping. Everything else can be revisited later.
Overview
The analog/digital boundary sits between the anti-alias filter and the sampler — and that placement is the whole point.
The sensors you will meet
| Sensor | Typical rate | Bits | Interface | Active power |
|---|---|---|---|---|
| MEMS microphone | 16–48 kHz | 16–24 | PDM / I²S | ≈ 0.3–1 mW |
| Accelerometer | 10 Hz – 6.4 kHz | 12–16 | I²C / SPI | ≈ 5–200 µW |
| Gyroscope | 25 Hz – 6.4 kHz | 16 | I²C / SPI | ≈ 1–3 mW |
| PPG (optical HR) | 25–500 Hz | 16–20 | I²C | ≈ 0.1–2 mW |
| ECG front end | 125–1000 Hz | 12–24 | SPI / ADC | ≈ 10–300 µW |
| Low-res camera | 1–30 fps | 8 / pixel | DCMI / MIPI | ≈ 1–100 mW |
The gyroscope costs 10–100× the accelerometer. And the sensor's configuration is part of the model's input spec.
01
A primer — the language every later step is described in.
Fourier
Primer · the DFT
Each X[k] measures how strongly the block correlates with a sinusoid at f_k. A longer block resolves finer frequency differences — at the price of time resolution.
Primer · convolution
01
Stated precisely, then demonstrated numerically.
The theorem, and what it really says
Sampling in time multiplies by an impulse train, which convolves the spectrum with an impulse train. The spectrum is periodically replicated. If f_s < 2B the replicas overlap and ADD — and an addition whose terms were never observed apart cannot be undone.
Mechanism
Replicas overlap → the observed spectrum is their sum.
Computed, not asserted
A 7 Hz tone sampled at 10 Hz produces exactly the sample set a 3 Hz tone would produce.
02
Where the second irreversible loss happens — and how much of it you are allowed.
Derivation
Assumption: the rounding error is uniform on [−q/2, q/2] and independent of the signal. Good when the signal is busy relative to q — and false when it is not.
What a quantizer really does
Every flat tread is a set of real inputs the converter cannot tell apart.
Design against this, not the label
ENOB = (SINAD − 1.76) / 6.02. A “16-bit” ADC on a noisy board often measures 11.
Foundations
| Architecture | Mechanism | Typical | Implication for you |
|---|---|---|---|
| SAR | binary search, one comparison per bit | 8–18 bit kSPS–MSPS | The MCU default. Low latency, easy to multiplex across channels. |
| Sigma–delta | 1-bit modulator at 64–256× oversampling + digital decimation | 16–24 bit kSPS | Audio and precision sensing. Adds group delay — matters for event timestamps and sensor fusion. |
| Flash / pipeline | 2ᴺ−1 parallel comparators, or a pipeline of stages | 6–14 bit MSPS–GSPS | Radar, ultrasound, RF. You will be bandwidth-limited downstream long before the converter is the problem. |
Each buys resolution with a different currency: comparisons, time, or silicon area.
03
What actually limits your SNR — and it is rarely the quantizer.
Choosing a structure
| Property | FIR | IIR |
|---|---|---|
| Cost per output sample | M MACs (M taps) | ≈ 2 × order (biquads) |
| State memory | M words | 2 words per biquad |
| Phase | exactly linear if symmetric | non-linear; group delay varies |
| Stability | always stable | poles must stay inside the unit circle |
| Fixed-point failure | accumulator overflow | coefficient quantization moves poles; limit cycles |
| Use for | decimation, matched filtering, anything phase-sensitive | cheap band-limiting, DC blocking, envelopes |
For a given magnitude spec an IIR is typically an order of magnitude cheaper. You pay in phase linearity and numerical fragility.
Two cheap filters
Sampling rate
Keyword spotting: 16 kHz conventional, 8 kHz often enough. Bearing faults: the diagnostic energy may sit above 10 kHz and the same reasoning gives the opposite answer.
Break
Five minutes.
Next: datasets and preprocessing, then features.
02
The most expensive and least reusable part of most projects.
Building a dataset
Normalisation
The commonest invalid result
Split by subject, device or recording — never by window.
Overlapping windows from one recording are near-copies. A random split measures memorisation.
Bridge to Session 2
On an energy-limited device, overlap is a parameter to be justified, never a default.
Segmentation
Method
If the front end costs 40 % of the budget, a learned front end on raw samples may be cheaper overall.
Time-domain features
01
Windows, leakage, and a bound you cannot tune your way around.
Definition
The Fourier transform assumes stationarity over its whole support. Speech, vibration and gesture signals are not stationary — so we apply it to short segments in which stationarity is approximately true.
Measured, not asserted
Rectangular −13 dB · Hann −31 · Hamming −43 · Blackman–Harris −92.
The bound
02
Four steps, each of which had a reason — and one of which no longer applies.
Perceptual warping
m(f) = 2595 · log₁₀(1 + f/700). Six filters below 1 kHz; five across the whole 4–8 kHz span.
The pipeline, and the branch that matters
The DCT existed to decorrelate for diagonal-covariance GMMs. A CNN has no such requirement — and the DCT destroys the local structure convolutions exploit.
The same sound, two representations
Voiced harmonics, a fricative burst, a click — all still legible after mel compression.
Price the front end
| Front end | Cost / frame | Output dim | Note |
|---|---|---|---|
| Time-domain statistics, 3-axis IMU, 128 samples | ~10³ ops | ~30 | Runs in the sensor ISR |
| 512-point real FFT | ~10⁴ MACs | 257 | CMSIS-DSP radix-2/4 |
| 40-band log-mel from that FFT | ~1.5 × 10⁴ | 40 | Filterbank is sparse |
| 13 MFCC + Δ + ΔΔ | ~1.6 × 10⁴ | 39 | Adds a small DCT |
| Small DS-CNN over a 49×10 map | ~2.7 × 10⁶ | — | Per inference ≈ 4× the front end of a whole second (49 frames ≈ 7 × 10⁵) |
The comparison flips for very small models and high sample rates. Measure, do not assume.
03
Which features survive — and the hygiene rule that decides whether a result is valid.
Three families
Filter — score each feature alone
Wrapper & embedded
The hygiene rule
Feature selection is part of the model, so it happens inside each cross-validation fold.
Selecting on the full dataset and then cross-validating the classifier leaks label information — routinely several accuracy points, and far more when features outnumber samples. The same applies to normalisation statistics, PCA bases and class balancing.
Dimensionality reduction
Supervised projection
Looking at data
Before Session 3
Y. Zhang, N. Suda, L. Lai, V. Chandra, “Hello Edge: Keyword Spotting on Microcontrollers,” 2017. arXiv:1711.07128.
Why: A complete embedded pipeline — MFCC front end, several model families, memory and operation budgets — in twelve pages. Session 3 uses its numbers.
No ML background? Work through Part A of the Session 3 web chapter before the lecture.
If you remember one thing
Decide what the model must distinguish, then work backwards: bandwidth sets the sampling rate, dynamic range sets the bits, event duration sets the window, and the cheapest representation that keeps the distinguishing information is the right one.
Before you go