Embedded Machine LearningPhD course

Session 2 · 90 minutes · lecture

Data acquisition, preprocessing and feature engineering

How a physical quantity becomes numbers, what is lost on the way, and how those numbers become a representation a small model can use. Sampling, quantization and noise; datasets and preprocessing; windows, spectrograms, MFCCs, IMU features; selection and dimensionality reduction.

Duration90 min, with a short break
PrerequisitesSession 1; complex numbers helpful, not required
SlidesPDF · HTML
Core readingWarden 2018; Zhang et al. 2017
Learning outcomes
  • Describe the sensing chain from transducer to digital samples, and name the loss introduced at each stage.
  • Explain the frequency view of a signal well enough to reason about bandwidth, filtering and sampling rate.
  • State the sampling theorem precisely, recognise aliasing, and specify an anti-aliasing filter and sampling rate for a task.
  • Compute quantization SQNR and ENOB, and read the noise specification of a sensor or converter.
  • Build a dataset responsibly: collection, labelling, cleaning, normalisation, and splits that do not leak.
  • Derive the STFT trade-off between time and frequency resolution, and compute log-mel and MFCC features step by step, with their cost.
  • Choose features for IMU and audio tasks, select among them, and reduce dimensionality with PCA or LDA — and know why t-SNE is not a deployment tool.
Timing plan · 90 minutes
  1. 0–6Recap of the five budgets; today's question: what has already been lost before the model sees the data?
  2. 6–16The sensing chain; the sensors you will meet; signals in time and frequency (Fig 2.2).
  3. 16–30Sampling and aliasing, with the interactive explorer; quantization, SQNR and ENOB.
  4. 30–40Noise and filtering (moving average versus median); sampling rate and resolution as energy decisions.
  5. 40–45Datasets, labels, normalisation, and leakage.
  6. 45–52Break.
  7. 52–58Windows and segmentation (Fig 2.9); three questions to ask of any representation.
  8. 58–74Time-domain features; the STFT and its resolution trade-off; mel and MFCC derived on the board.
  9. 74–84IMU and image features; feature selection; PCA, LDA, and what t-SNE is for.
  10. 84–90Takeaway; preview of Session 3.
In one sentence

Every stage between the physical world and the model — transducer, filter, sampler, quantizer, window, transform — throws information away, and the engineering question is always whether it threw away the part the model needed, at a price the budget can pay.

2.1The chain, and what each stage costs

transducer+ bias amplifier/ AGC anti-aliasfilter sample& hold quantizerN bits digital filter+ decimate frame, window,features continuous time — analog domain discrete time — digital domain thermal +1/f noise clipping,distortion band abovef_s/2 (wanted) jitter →phase noise q²/12noise power transitionband ripple time/frequencyresolution every red annotation is information the model can never recover
Fig 2.1The sensing chain. The boundary between analog and digital sits between the anti-alias filter and the sampler — a filter placed after sampling cannot undo aliasing, because folding has already made the aliased and wanted components indistinguishable.

Analog versus digital sensors. An analog sensor delivers a continuous voltage and requires the whole chain above; you own the anti-alias filter, the reference, and the noise budget. A digital sensor (an I²S MEMS microphone, an I²C IMU) has the converter on-die and hands you samples over a bus — convenient, but the conversion decisions have been made for you and are visible only in the datasheet. The important point is that "digital sensor" does not mean "no signal processing"; it means the processing happened somewhere you cannot change, so you must read its specification: sensitivity, acoustic overload point, SNR, and the on-chip decimation filter's group delay.

2.2The sensors you will meet

A sensor (more precisely, a transducer plus its conditioning electronics) converts a physical quantity — pressure, acceleration, light, voltage across skin — into an electrical signal. Most sensors used in embedded ML today are MEMS devices (micro-electro-mechanical systems): microscopic mechanical structures etched in silicon, packaged with their own amplifier and often their own analog-to-digital converter, which deliver digital samples over a serial bus. For the ML engineer, four properties of a sensor matter more than its physics: its sampling rate, its resolution and noise floor, its interface, and its power draw while sampling.

SensorMeasuresTypical rateBitsInterfaceActive power
MEMS microphonesound pressure16–48 kHz16–24PDM or I²S≈ 0.3–1 mW
Accelerometerlinear acceleration, 3 axes10 Hz–6.4 kHz12–16I²C / SPI≈ 5–200 µW
Gyroscopeangular rate, 3 axes25 Hz–6.4 kHz16I²C / SPI≈ 1–3 mW
PPG (optical heart rate)blood-volume changes25–500 Hz16–20I²C≈ 0.1–2 mW (LED)
ECG front endcardiac potential125–1000 Hz12–24SPI / ADC≈ 10–300 µW
Environmentaltemperature, humidity, gas0.01–10 Hz12–24I²CµW (gas: mW, heated)
Low-res cameralight intensity, 2-D1–30 fps8 per pixelDCMI / SPI / MIPI≈ 1–100 mW
mmWave radarrange, velocity, angleframes at 10–100 Hz12–16SPI / LVDS≈ 0.1–1 W (duty-cycled)

Table 2.1 — Sensors in embedded ML. Order-of-magnitude ranges across common parts; the active-power column is the sensor alone. Note that a gyroscope costs one to two orders of magnitude more power than an accelerometer — a classic reason to infer orientation from the accelerometer alone when the application allows it.

A three-axis MEMS accelerometer breakout board. The sensor itself is the 3 × 3 mm package in the centre.
A three-axis MEMS accelerometer breakout board. The sensor itself is the 3 × 3 mm package in the centre.

Photo: Wikimedia Commons

A common 6-axis IMU module (accelerometer and gyroscope) used in countless student and research projects.
A common 6-axis IMU module (accelerometer and gyroscope) used in countless student and research projects.

Photo: Wikimedia Commons

Schematic of an electret condenser microphone. MEMS microphones work on the same capacitive principle, scaled down onto silicon.
Schematic of an electret condenser microphone. MEMS microphones work on the same capacitive principle, scaled down onto silicon.

Photo: Wikimedia Commons

Two practical lessons recur. First, many sensors can do part of the processing themselves: modern IMUs contain FIFO buffers, step counters and even small programmable state machines or decision-tree engines, so the main MCU can sleep until something interesting happens — the cascade of Session 1 implemented inside the sensor. Second, the sensor's configuration (range, rate, filter) is part of the model's input specification: a classifier trained on data captured at ±4 g and 100 Hz will fail silently on a device configured for ±16 g and 50 Hz. Record the configuration with the data.

2.3Signals in time and frequency — a primer

A signal here is a sequence of numbers indexed by time, x[n] = x(n·Ts), where Ts is the sampling interval and fs = 1/Ts the sampling rate. You can describe it in two equivalent ways: by its values over time, or by how much of each frequency it contains. The second description is the key to almost everything in this session.

The bridge is Fourier's observation that any reasonable signal is a sum of sinusoids. A sinusoid A·sin(2π f t + φ) has three properties: amplitude A, frequency f (cycles per second, Hz) and phase φ. The spectrum of a signal lists the amplitude (and phase) of each frequency it contains.

0 T/2 T 3T/2 2T −1 0 1 time square wave = sum of odd sine harmonics partial sums — grey: 1 harmonic · teal: 3 · copper: 10 the ~9 % overshoot at the jumps never vanishes (Gibbs) f₀ 3f₀ 5f₀ 7f₀ 9f₀ 11f₀ 0 0.5 1.0 frequency its spectrum: amplitude 4/(πk) sharp edges = slowly decaying harmonics = bandwidth to handle
Fig 2.2Fourier's idea in one picture. Any periodic signal is a sum of sinusoids; a square wave needs every odd harmonic, with amplitudes falling only as 1/k. The practical lesson for sensing is on the right: the sharper the features of a signal in time, the further its energy extends in frequency, and the higher the sampling rate (or the stronger the anti-aliasing filter) it demands.
Foundations · the discrete Fourier transform in five lines

For a block of N samples, the discrete Fourier transform (DFT) computes N complex numbers

X[k] = Σₙ x[n] · e^(−j 2π k n / N), k = 0 … N−1

Each X[k] measures how strongly the block correlates with a sinusoid at frequency fk = k · fs/N; its magnitude |X[k]| is the amplitude, its angle the phase. Three facts follow directly. The frequency resolution is fs/N: a longer block resolves finer frequency differences. For a real signal only the first N/2 + 1 bins are independent; the highest representable frequency is fs/2. And the fast Fourier transform (FFT) computes all N bins in about N·log₂N operations instead of N² — for N = 512, on the order of 4 600 complex operations (2 304 radix-2 butterflies) instead of N² ≈ 262 000 complex multiplications, which is why spectral features are affordable on a microcontroller at all.

Foundations · convolution, filtering and the frequency domain

A linear, time-invariant filter is fully described by its impulse response h[n]: its output is the convolution y[n] = Σₖ h[k]·x[n−k], a sliding weighted sum — the moving average is the simplest example. The convolution theorem says that convolution in time is multiplication in frequency: Y(f) = H(f)·X(f). A low-pass filter is therefore a multiplication of the spectrum by a function that is near 1 at low frequencies and near 0 at high ones. The theorem also works the other way round: multiplying two signals in time convolves their spectra. Sampling is exactly such a multiplication — the continuous signal times a train of impulses spaced Ts apart — and the spectrum of an impulse train is itself an impulse train spaced fs apart. Convolving with it makes copies of the signal's spectrum at every multiple of fs, which is the picture behind the sampling theorem and aliasing in the next section. (The same word, convolution, names the core layer of CNNs in Session 3: there too it is a sliding weighted sum, with learned weights.)

Why should an ML engineer care? Because most physical processes are easier to recognise in frequency than in time. A bearing fault produces vibration at a characteristic frequency; a spoken vowel is a pattern of resonances; walking is a periodic motion at about 2 Hz. And because the operations of the sensing chain — filtering, sampling, windowing — have simple descriptions in the frequency domain and complicated ones in time.

2.4Sampling and aliasing, stated precisely

Sampling theorem. A signal whose spectrum is zero outside |f| < B is completely determined by samples taken at rate fs > 2B, and is reconstructed by sinc interpolation. Sampling in time multiplies by an impulse train, which convolves the spectrum with an impulse train of spacing fs: the spectrum is periodically replicated. If fs < 2B, adjacent replicas overlap and add. That overlap is aliasing, and it is an addition, not an occlusion — which is why it is irreversible.

A component at frequency f appears after sampling at the folded frequency

f_alias = | f − round(f / f_s) · f_s |

so a 30 kHz interferer sampled at 16 kHz lands at 2 kHz, squarely inside a speech band, and is now indistinguishable from real speech energy.

f_s > 2B — replicas separated, baseband recoverable baseband f_s 2f_s guard band f_s < 2B — replicas overlap; the sum in the shaded band cannot be undone aliased energy
Fig 2.3Aliasing as spectral folding. Sampling replicates the spectrum at multiples of fs. When the replicas overlap, the observed spectrum is their sum; no digital filter can separate a sum whose terms were never observed apart. Hence the anti-alias filter must be analog and must precede the sampler.
7 Hz 3 Hz alias true 7 Hz · f_s = 10 Hz · Nyquist limit 5 Hz teal dots = everything the converter ever sees | 7 − round(7/10)·10 | = 3 Hz Same samples, two different signals — the loss is already permanent.
Fig 2.4Aliasing, computed rather than asserted. A 7 Hz tone sampled at 10 Hz produces exactly the sample set a 3 Hz tone would produce. The two curves are not confusable; after sampling they are identical, so no amount of digital filtering can separate them. This is why the anti-alias filter must be analog and must sit before the sampler — it has to remove the 7 Hz energy while the two signals are still different things.
InteractiveAliasing explorer — push fs below 2f and watch the copper curve appear
Nyquist limit–
apparent frequency–
…
Design consequence students routinely miss

A brick-wall analog filter does not exist. Real filters have a transition band, so practical designs sample at 2.2–2.5× the highest wanted frequency and let the filter roll off across the guard band, or oversample heavily and do the sharp filtering digitally (§2.7). "Sample at exactly 2B" is a theorem, not a design.

2.5Quantization noise, SQNR and ENOB depth

Derivation — the 6.02 N + 1.76 dB rule

An N-bit converter with full-scale range VFS has step size q = VFS/2N. Model the rounding error e as uniformly distributed on [−q/2, q/2] and independent of the signal — a good approximation when the signal is busy relative to q. Its power is

σ_e² = (1/q) ∫ from −q/2 to q/2 of e² de = q² / 12

A full-scale sinusoid of amplitude VFS/2 has power VFS²/8. Therefore

SQNR = 10 log₁₀ ( (V_FS²/8) / (q²/12) ) = 10 log₁₀ ( 1.5 · 2^(2N) ) ≈ 6.02 N + 1.76 [dB]

Each additional bit buys about 6 dB. Two corollaries matter in practice:

  • Headroom is expensive. If your signal only ever reaches −20 dBFS, you have thrown away more than three bits. Gain staging is therefore part of the ML pipeline, not merely electronics.
  • Nominal bits ≠ effective bits. Datasheets quote ENOB, derived from measured SINAD by inverting the formula: ENOB = (SINAD − 1.76)/6.02. A "16-bit" ADC on a noisy PCB may deliver 11–12 ENOB. Design against ENOB.

Dither. The independence assumption fails for very small or very periodic signals, where the error becomes correlated with the signal and appears as harmonic distortion — audible as tones rather than hiss. Adding a small amount of noise (about 1 LSB rms) before quantization decorrelates the error, trading a slightly higher noise floor for the removal of structured artefacts. This matters for embedded audio because structured artefacts are exactly what a spectrogram-based classifier will latch onto.

-1 0 1 -1 0 1 3-bit quantizer · 8 levels ideal input error e = x̂ − x +q/2 −q/2 σ²ₑ = q²/12 → SQNR ≈ 6.02 N + 1.76 dB Every flat tread is a set of inputs the converter cannot tell apart.
Fig 2.5What a quantizer actually does. The error is not noise — it is a deterministic sawtooth function of the input. We are allowed to model it as uniform white noise only when the signal moves across many steps between samples. For a quiet or strongly periodic signal that assumption fails: the sawtooth becomes correlated with the signal and appears as harmonic distortion, which a spectrogram-based classifier will happily learn as a feature. That failure mode is exactly what dither exists to prevent.
0 4 8 12 16 20 24 0 40 80 120 160 int8 · 50 dB 12-bit · 74 dB 16-bit · 98 dB 24-bit · 146 dB a “16-bit” ADC on a noisy board often measures 11 ENOB → 30 dB gone before the model sees anything nominal bits, N SQNR (dB) SQNR = 6.02 N + 1.76 dB Six decibels per bit — and six decibels of unused headroom throws one away.
Fig 2.6Bits are worth 6 dB each, in both directions. The line is the theoretical best; the copper marker is what a real board delivers. Two practical consequences: quote and design against ENOB rather than the nominal word length, and treat gain staging as part of the machine-learning pipeline — a signal that only ever reaches −20 dBFS has already thrown away more than three bits before any model is trained.

Averaging and oversampling. Averaging M independent samples of a stationary signal reduces noise power by M, i.e. SNR improves by 10 log₁₀ M dB — about 3 dB per doubling, or half a bit. Getting one extra effective bit therefore costs 4× the samples. This is the exchange rate between analog quality and digital effort, and it is the basis of sigma-delta converters, which deliberately oversample by 64–256× and shape the quantization noise out of the band of interest.

2.6Noise budget and filtering

Sources, in the order they usually dominate on a real board: power-supply and switching noise coupled through the reference; electromagnetic interference at mains frequency and its harmonics; thermal (Johnson) noise, white, proportional to √(4kTRB); 1/f (flicker) noise, dominant at low frequencies and the reason DC-coupled measurements are hard; quantization noise, usually the smallest of the five. Note the ordering: students are taught quantization noise and then find that it is rarely the limiting term.

Foundations · three converter architectures, and what each implies for you
ArchitectureHow it worksTypicalConsequence for an ML pipeline
SAR (successive approximation)Binary search: one comparison per bit against a DAC8–18 bit,
kSPS–MSPS
The MCU default. One conversion per trigger, low latency, easy to interleave across channels — good for IMU and general sensing.
Sigma–delta (ΣΔ)1-bit modulator at 64–256× oversampling, then a digital decimation filter shapes the noise out of band16–24 bit,
kSPS
Audio and precision sensing. Buys resolution with digital effort instead of analog precision — but its decimation filter adds group delay, which matters if you are timestamping events or fusing sensors.
Flash / pipeline2N−1 comparators in parallel, or a pipeline of small stages6–14 bit,
MSPS–GSPS
Radar, ultrasound, RF. Fast and power-hungry; you will be bandwidth-limited downstream long before the converter is the problem.

Table 2.2 — Converter architectures. Notice the pattern: each buys resolution with a different currency — comparisons (SAR), time and digital filtering (ΣΔ), or silicon area and power (flash). That is the same trade structure you will meet again in Session 4, one abstraction level up.

Foundations · the decibel, and why everything here is logarithmic

A decibel is a ratio of powers on a log scale: 10 log₁₀(P₁/P₀), or equivalently 20 log₁₀(A₁/A₀) for amplitudes, since power goes as amplitude squared. Three numbers are worth internalising:

  • 3 dB = a factor of 2 in power. Averaging two independent samples, or doubling the oversampling ratio, buys you exactly this.
  • 6 dB = a factor of 2 in amplitude — one bit. Hence §2.5's rule.
  • 20 dB = a factor of 10 in amplitude, 100 in power.

The reason the whole signal chain is measured this way is that its stages multiply: a sensor's sensitivity, an amplifier's gain and a converter's full-scale range compose by multiplication, and logarithms turn that into addition. It is also how a noise budget is assembled — powers of independent noise sources add, so the total is √(σ₁² + σ₂² + …) and one source 10 dB above the rest simply is the answer.

Watch the reference, which is where mistakes are most common: dBFS is relative to the converter's full scale and is therefore never positive; dBSPL is relative to 20 µPa of acoustic pressure; dBA is dBSPL after a perceptual weighting curve. A microphone datasheet quoting “65 dB SNR” means SNR at 94 dBSPL, A-weighted — a different quantity from the SQNR of the converter behind it.

FIR versus IIR on an embedded target

PropertyFIRIIR
Cost per output sampleM MACs (M taps)≈ 2·order MACs (biquads)
State memoryM words2 words per biquad
PhaseExactly linear if symmetricNon-linear; group delay varies
StabilityAlways stableCan oscillate; poles must stay inside the unit circle
Fixed-point pathologiesOverflow in the accumulatorCoefficient quantization moves poles; limit cycles; needs cascaded biquads and careful scaling
Typical useDecimation, matched filtering, anything where phase mattersCheap band-limiting, DC blocking, envelope tracking

Table 2.3 — Choosing a filter structure. For a given magnitude specification an IIR filter is typically an order of magnitude cheaper than an FIR one; you pay in phase linearity and numerical fragility.

The moving average is an FIR filter with impulse response 1/M over M taps, and its frequency response is a Dirichlet (periodic sinc) kernel: a wide main lobe and first sidelobe only about 13 dB down. It is cheap and it is a poor low-pass filter, and it is worth looking at its magnitude response once, to stop treating "smoothing" as free. The moving median, by contrast, is non-linear: it removes impulsive outliers without smearing edges, has no frequency response at all, and costs a sort or a running-order-statistic structure — often the right choice for IMU spike removal and the wrong choice for anything you will later transform.

0 1 moving average, 9 taps — linear, cheap, smears spikes and edges 0 1 median, 9 taps — non-linear, removes spikes, keeps the edge grey: raw samples (step + Gaussian noise + six impulsive spikes)
Fig 2.7Two of the cheapest filters on a microcontroller, applied to the same signal. The moving average is a low-pass FIR filter: it reduces Gaussian noise well but turns each spike into a bump and blurs the step. The median filter is non-linear: it rejects outliers completely and preserves edges, at the price of a small sort per sample. Which one is "right" depends on whether the spikes are noise (a loose connector) or signal (a knock you want to detect).

Fixed-point implementation. On MCUs without an FPU, filters run in Q-format fixed point. Three rules: accumulate in a wider type than you multiply in (Q15 × Q15 → Q30 in a 32-bit accumulator); scale coefficients so intermediate results cannot overflow, or use saturating arithmetic; and implement high-order IIR filters as cascaded second-order sections, because direct-form high-order implementations are numerically catastrophic. CMSIS-DSP's biquad functions exist for exactly these reasons.

2.7Sampling rate and resolution are energy decisions

Sampling rate propagates multiplicatively through the whole downstream cost: ADC energy per conversion × rate, filter MACs per sample × rate, frames per second × feature cost, inferences per second × model cost. Halving fs where the application allows it is usually the single largest energy saving available, and it is free of accuracy loss if the discarded band carries no class-discriminative information — a claim that should be tested empirically, not assumed. For keyword spotting, 16 kHz is conventional and 8 kHz is often sufficient; for bearing-fault detection the diagnostic energy may sit above 10 kHz and the same reasoning gives the opposite answer.

Oversample-and-decimate is the standard escape from expensive analog design: sample at R·fs with a gentle analog filter, apply a sharp digital low-pass, then keep every R-th sample. Two efficiency notes for embedded targets: the decimating filter should be implemented as a polyphase structure so it only computes the outputs that survive (an R× saving), and cascaded integrator-comb (CIC) filters achieve large decimation ratios with no multipliers at all, which is why they appear inside PDM microphone front ends.

2.8Datasets, labels and preprocessing

The signal chain delivers samples; a learning algorithm needs a dataset: many examples, each paired with a label, that represent the conditions the device will meet in the field. For embedded sensing problems the dataset is usually the most expensive and least reusable part of the project, and the most common cause of disappointing field performance.

Collection and labelling

  • Coverage beats volume. Record across the variation the device will see: different users, mounting positions, devices of the same model (sensor gains differ), environments, background activities. A thousand recordings from one person are worth less than a hundred from fifty.
  • Label granularity is a design decision. A recording may be labelled as a whole ("this minute contains a fall"), per window, or with exact event boundaries. Coarse labels are cheap; fine labels are needed to train detectors with accurate onsets. Weakly labelled learning is an active research area precisely because fine labels are so expensive.
  • Rare events and class imbalance. The events you care about — faults, falls, arrhythmias, a wake word — are usually rare. Plan for it: collect targeted examples, use a background ("none of the above") class that is genuinely diverse, and evaluate with metrics that are not fooled by imbalance (Session 3).
  • Public datasets. Speech Commands (keyword spotting), ESC-50 and UrbanSound8K (environmental audio), UCI-HAR and PAMAP2 (activity recognition), MIT-BIH (arrhythmia), MIMII and ToyADMOS (machine sound anomalies), Visual Wake Words (person detection). They are ideal for comparison and teaching, and dangerous as a proxy for your own deployment conditions.

Cleaning and preprocessing

Before features, the raw data usually needs: synchronisation of streams sampled by different clocks; resampling to the rate the deployed device will use; handling of gaps (dropped packets, sensor saturation) — by interpolation, masking or discarding windows; outlier and artefact removal (the median filter above); and calibration (offset and gain per sensor). Every one of these steps must also exist, identically, in the firmware — a discrepancy between the training pipeline in Python and the deployed pipeline in C is one of the most common causes of embedded models that "worked in the notebook".

Normalisation

raw values RMS level, mV zero-crossing rate spectral peak, Hz 0 1000 2000 z-score: (x − μ)/σ −2 0 2 6 min-max: (x − min)/(max − min) 0 .5 1 Raw: features 1 and 2 vanish next to feature 3 — a distance-based model (k-NN, RBF-SVM, k-means) effectively ignores them. Min-max: one outlier (copper) squeezes the other 59 values of feature 3 into a fifth of the range. Fit all scaling statistics on the training set only.
Fig 2.8Why features are scaled before most models see them. Raw, the feature with the largest numerical range dominates every Euclidean distance and every gradient step. Z-scoring gives each feature zero mean and unit variance; min-max scaling maps to [0, 1] but lets a single outlier (copper) dictate the scale. On a device the scaling constants are fixed numbers baked into the firmware, computed on the training set only — computing them on the test set is one of the commonest forms of data leakage.

Most learning algorithms assume features on comparable scales. Standardisation (z-scoring) subtracts the training-set mean and divides by the training-set standard deviation of each feature; min–max scaling maps the training range to [0, 1]. Decision trees are indifferent to scaling; neural networks, SVMs, k-nearest neighbours and anything trained by gradient descent are not (the reason is visible in the gradient-descent figure of Session 3). On an embedded device the normalisation is two constants per feature, applied in the firmware — and, conveniently, it can be folded into the first layer's weights and the quantization parameters (Session 3), so it often costs nothing at run time.

Augmentation

Data augmentation creates plausible new training examples from existing ones: time shifts, added background noise at various signal-to-noise ratios, gain changes and room reverberation for audio; small rotations of the sensor frame, time warping and amplitude scaling for IMU data. It is free at inference time and often the cheapest accuracy improvement available to a small model, because small models are more sensitive to the gap between training and field conditions. SpecAugment-style masking of time and frequency bands is standard for spectrogram inputs.

Common misconception

"We split the data randomly into training and test sets, so the test accuracy is honest." Not for windowed sensor data. Overlapping windows cut from the same recording are near-copies; a random split puts copies on both sides, and the test score measures memorisation of that recording. The same holds, less obviously, for windows from the same person or the same machine. Split by subject, device or session — Session 3 shows the picture. Fit normalisation constants and feature selection on the training portion only, for the same reason.

2.9Framing, windowing and segmentation

Models consume frames, not streams. A frame length L with hop H gives an overlap of 1 − H/L and a frame rate of fs/H. Three consequences: the frame rate sets the inference rate and therefore the energy; the frame length sets the lowest frequency that can be resolved (fs/L); and the hop sets the temporal precision of an event's onset. Overlap is not free — 50 % overlap doubles the number of inferences — so on an energy-limited device it is a parameter to be justified, not a default. §2.12 takes this up properly with the time–frequency uncertainty relation.

0 s 2 4 6 8 10 12 s −1 g 0 1 2 g standing walking running windows: 2 s (100 samples) with 50 % overlap → one feature vector every second σ(|a|) one simple feature per window — the standard deviation of |a| — already separates the three activities
Fig 2.9From a stream to examples. A wearable's accelerometer (copper x, teal y, grey z, which carries gravity) is cut into overlapping windows; each window becomes one feature vector and one classification. Even a single feature — the standard deviation of the acceleration magnitude — separates standing, walking and running here. Window length trades latency against statistical stability, and overlap trades compute against temporal resolution. Synthetic signal for illustration.

For event-like data the fixed sliding window is not the only option. Event-triggered segmentation starts a window when a cheap detector (an energy threshold, a step detector in the IMU) fires, which aligns windows with the events and saves energy between them. Labels must be assigned per window — by majority vote over the samples in it, by "any event present", or by the label at the window centre — and that choice changes what the model learns about onsets. Write it down.

2.10Three questions to ask of any representation

  1. What does it discard, and is the discarded part class-discriminative? Mel warping discards fine high-frequency resolution — irrelevant for speech, potentially fatal for ultrasonic fault detection.
  2. What invariances does it build in? The magnitude spectrogram discards phase and thereby buys invariance to time shift within a frame. If your task depends on inter-channel phase (direction of arrival), you have just deleted the signal.
  3. What does it cost, and against what baseline? Compare the front end's MACs and memory with the model's. If the front end costs 40 % of the budget, a learned front end operating on raw samples may be cheaper overall.

2.11Time-domain features

The cheapest features are computed directly on the samples of a window, with one pass and no transform. For a window x[0…N−1]:

mean μ = (1/N) Σ x[n] energy / RMS E = √( (1/N) Σ x[n]² ) std. dev. σ = √( (1/N) Σ (x[n] − μ)² ) zero-crossing ZCR = (1/(N−1)) Σ 𝟙[ sign x[n] ≠ sign x[n−1] ] peak-to-peak max x − min x

plus higher-order statistics (skewness, kurtosis), percentiles, and counts of threshold crossings. Each costs O(N) operations — tens of microseconds on an MCU for a typical window — and many can be updated incrementally as each sample arrives, so the window never needs to be stored at all.

voiced vowel fricative /s/ silence RMS energy zero- crossing rate frames: 20 ms, hop 10 ms · both features cost one pass over the frame, no FFT
Fig 2.10Two time-domain audio features that cost almost nothing. Short-time energy (copper) finds where the sound is; zero-crossing rate (teal) tells voiced sounds, whose energy is at low frequencies, from noisy fricatives, whose energy is at high frequencies. Silence also has a high zero-crossing rate, because low-level noise crosses zero constantly, which is why neither feature is used alone. Together they make a classic voice-activity detector — exactly the kind of always-on first stage the cascade of Session 1 calls for. Synthetic signal for illustration.

These features are what the first stage of a cascade is built from: a voice-activity detector, a "something is moving" trigger for an IMU, or an "unusual vibration energy" alarm. They are also more powerful than they look for slowly varying physical signals, and a sensible baseline for any new task: if a decision tree on ten time-domain statistics already reaches the target accuracy, the project does not need a neural network.

2.12The short-time Fourier transform depth

The Fourier transform assumes stationarity over its whole support; speech, vibration and gesture signals are not stationary. The STFT applies the transform to short windowed segments in which stationarity is approximately true:

X[m, k] = Σ over n of x[n] · w[n − mH] · exp(−j2πkn/L)

with window w of length L and hop H. The spectrogram is |X[m,k]|².

Windows and leakage

Truncating with a rectangular window multiplies in time, hence convolves in frequency with a Dirichlet kernel whose first sidelobe is only ≈13 dB down. A strong component therefore leaks into neighbouring bins and can bury a weak one. Tapered windows trade main-lobe width (resolution) for sidelobe suppression (dynamic range):

WindowMain lobe widthPeak sidelobeUse when
Rectangular2 bins−13 dBTransient onsets; never for spectral estimation
Hann4 bins−31 dBDefault for speech/audio; sums to unity at 50 % overlap
Hamming4 bins−43 dBClassic speech front ends
Blackman–Harris8 bins−92 dBLarge dynamic range, e.g. vibration with a dominant shaft tone

Table 2.4 — Window trade-offs. Widths are approximate, in units of the DFT bin spacing fs/L. Choosing a window is choosing where on the resolution/dynamic-range curve your task sits.

0 2 4 6 8 0 -20 -40 -60 -80 -100 rectangular · -13 dB Hann · -31 dB Hamming · -43 dB Blackman–Harris · -92 dB frequency offset, in DFT bins magnitude (dB) peak sidelobe level, measured by DFT of a 64-point window Narrow main lobe or low sidelobes — you cannot have both.
Fig 2.11The window trade, measured. A narrow main lobe lets you resolve two tones that are close in frequency; low sidelobes stop a loud tone from burying a quiet one several bins away. The rectangular window wins the first contest and loses the second badly. Choose by asking which failure your task can tolerate: for speech, Hann or Hamming; for vibration analysis with a dominant shaft tone 60 dB above the fault signature, Blackman–Harris.

The uncertainty relation, and what it means operationally

Frequency resolution is Δf ≈ fs/L and time resolution is Δt = L/fs, so Δt·Δf ≈ 1 — a bound, not a tuning knob (for a Gaussian window the exact statement is σtσf ≥ 1/4π). You cannot resolve a 20 ms transient and a 10 Hz frequency difference in the same analysis. Increasing overlap increases the frame rate and thus the smoothness of the picture, but it does not improve Δt; the smearing is set by L. Students conflate these constantly.

short window L = 128 · Δt small, Δf large time frequency long window L = 1024 · Δt large, Δf small time every tile has the same area — the window chooses its shape, not its size
Fig 2.12Time–frequency tiling. The shaded tile is one STFT coefficient. Shortening the window makes tiles tall and narrow (good onset localisation, poor frequency discrimination); lengthening it does the opposite. Overlap changes how densely the plane is sampled, not the tile size.
InteractiveWindow length — the resolution trade, priced
frequency resolution–
time smearing–
frame rate–
front-end cost–
…

Front-end cost counts only the FFTs. Notice that overlap multiplies it without improving either resolution — the commonest wasted milliwatt in embedded audio.

2.13From spectrogram to mel and MFCC

Human auditory frequency resolution is approximately logarithmic above ~1 kHz. The mel scale encodes this; the most commonly used form is

m(f) = 2595 · log₁₀ ( 1 + f / 700 )
0 1k 2k 4k 6k 8k 0 1 26 triangular filters, equally spaced on the mel axis, drawn on a linear Hz axis 6 filters 5 filters over the same 4 kHz span frequency (Hz) m(f) = 2595 · log₁₀(1 + f / 700) Resolution is spent where human hearing has it — and thrown away where it does not.
Fig 2.13A mel filterbank is a deliberate, non-uniform loss of frequency resolution. Nine of these 26 filters have their centres below 1 kHz; only six lie above 4 kHz. For speech this is exactly right, because that is roughly how the cochlea allocates its own resolution. For a bearing-fault detector whose diagnostic energy sits at 6 kHz it is exactly wrong, and using a mel front end there is a mistake that no amount of model capacity will repair.

A mel filterbank places B triangular filters equally spaced on the mel axis and applies them to the power spectrum, reducing, say, 257 FFT bins to 40 mel bands — a 6× dimensionality reduction that is nearly free of task-relevant information for speech. The classical MFCC pipeline then continues:

pre-emphasis frame +window FFT power|·|² melfilterbank log DCT-II keep 1–13,lifter log-mel → stop here for a CNN cost per frame ≈ O(L log L) for the FFT + B·L/2 for the filterbank + B log B for the DCT → the filterbank is sparse in practice: each bin touches at most 2 triangles
Fig 2.14The MFCC pipeline, with the branch that matters. The DCT was introduced to decorrelate the log-mel energies so that Gaussian mixture models with diagonal covariance could be used. A convolutional network has no such requirement, and the DCT destroys the local frequency structure that convolutions exploit — which is why modern audio CNNs take log-mel spectrograms and stop before the DCT.
linear spectrogram · 257 bins × 138 frames mel spectrogram · 40 bands × 138 frames 8 kHz 4 kHz 0 band 40 band 20 band 1 time → time → voiced harmonics fricative noise click 6.4× fewer numbers, nearly all of the task-relevant structure.
Fig 2.15The same 1.2-second synthetic sound in two representations: a harmonic voiced segment with a moving pitch, a broadband fricative burst, and a click. The linear spectrogram spends half its 257 bins on the 4–8 kHz octave, where almost nothing distinguishes these three events; the 40-band mel version compresses that region and keeps the low-frequency harmonic structure intact. On a microcontroller this is a 6.4× reduction in the input tensor — and therefore in the first layer's cost — for a loss you can see is small.

Why each step exists. Pre-emphasis (y[n] = x[n] − αx[n−1], α ≈ 0.97) flattens the roughly −6 dB/octave spectral tilt of voiced speech so that high-frequency detail is not numerically swamped. Log compression converts multiplicative channel effects into additive offsets, which is what makes cepstral mean normalisation a valid channel-compensation method. DCT-II approximates the Karhunen–Loève transform for this class of signals, decorrelating the coefficients; keeping the first 12–13 discards the fine spectral detail that encodes pitch. Liftering rescales coefficients so their dynamic ranges are comparable.

Deltas. First and second time derivatives (Δ, ΔΔ), computed by regression over ±2 frames, add dynamic information to an otherwise static frame representation. For a CNN or an RNN with temporal context they are largely redundant, since the model can compute them; for an SVM or GMM over single frames they are essential. This is a good illustration of the general rule that feature engineering compensates for what the model class cannot learn.

2.14IMU and image features

IMU (accelerometer / gyroscope). A well-tested feature set for human-activity recognition, per axis per window: mean, standard deviation, min, max, interquartile range, root-mean-square, skewness, kurtosis, zero-crossing rate, signal-magnitude area, correlation between axes, and the norm √(ax²+ay²+az²), which is rotation-invariant and therefore robust to how the device is worn. Frequency-domain additions: dominant frequency, spectral entropy (flatness of the spectrum, distinguishing periodic walking from aperiodic fidgeting), band energies, and autocorrelation peak lag for cadence. Jerk (the derivative of acceleration) separates smooth from abrupt motions. Note that on a 50 Hz IMU a 2.56 s window is 128 samples — an FFT of that size is trivial, so the choice of time- versus frequency-domain features is about statistics, not cost.

Images. Handcrafted descriptors — edges, corners, colour histograms, HOG, LBP, Haar features over an integral image — were the state of the art before 2012 and remain competitive under extreme constraints, where the alternative model would have to be very small. In current practice a learned front end (a few convolutional layers of a MobileNet-class network) dominates whenever the compute exists, and the interesting question for this course is exactly where the crossover sits on a given device.

Front endApprox. cost / frameOutput dimNotes
Time-domain statistics, 3-axis IMU, 128 samples~10³ ops~30Negligible; runs in the sensor ISR
512-point real FFT~10⁴ MACs257CMSIS-DSP; radix-2/4
40-band log-mel from 512-point FFT~1.5 × 10⁴ MACs40Filterbank is sparse; log dominates if done naively
13 MFCC + Δ + ΔΔ~1.6 × 10⁴ MACs39Adds a small DCT
Small DS-CNN over a 49×10 MFCC map~2.7 × 10⁶ MACs—Model ≈ 4× the front end over a full second (49 frames ≈ 7 × 10⁵ MACs)

Table 2.5 — Order-of-magnitude front-end costs, for building the habit of comparing front end against model. The comparison flips for very small models and for high sample rates: measure, do not assume.

Common misconception

“Feature extraction is preprocessing, so it is free.” On a Cortex-M4F running a 40 kB keyword spotter at 16 kHz with 50 % overlap, the FFT-plus-mel front end can consume a third of the cycle budget and more than a third of the energy, because it runs on every frame while the classifier may run on a subsample of them. Profile the front end and the model separately, and count the front end's cost against the same budget as the model's.

2.15Feature selection depth

Three families, in increasing cost and decreasing generality:

  • Filter methods score each feature against the label independently of any model. Pearson correlation ρ = cov(X,Y)/(σXσY) captures only linear dependence and will assign zero to a perfect quadratic relationship. Mutual information, I(X;Y) = H(Y) − H(Y|X), captures any dependence but must be estimated — typically by k-nearest-neighbour estimators — and estimation error grows quickly with dimension. Filters are cheap and ignore redundancy: two perfectly correlated informative features both score highly and both get selected.
Foundations · entropy and mutual information in three lines

Entropy is the average number of nats (or bits, with log₂) needed to describe a draw from a distribution: H(Y) = −Σ p(y) log p(y). A fair coin has H = 1 bit; a coin that always lands heads has H = 0 — nothing to communicate.

Conditional entropy H(Y|X) is what remains uncertain about Y once you know X. Mutual information is the difference: I(X;Y) = H(Y) − H(Y|X) — literally how many bits knowing X saves you. It is zero exactly when the variables are independent, symmetric in its arguments, and invariant to any invertible transformation of either variable, which is why it does not care whether the relationship is a line, a parabola or a circle.

The price is estimation. Entropy of a continuous variable has to be estimated from samples — usually by binning or by nearest-neighbour distances — and the estimator's variance grows sharply with dimension. That is the honest reason Pearson correlation survives in practice despite being obviously weaker: it has a closed form and no hyperparameters.

linear Pearson ρ = +0.91 mutual info = 0.86 nats quadratic Pearson ρ = +0.09 mutual info = 1.02 nats circular Pearson ρ = +0.09 mutual info = 0.85 nats Pearson sees only straight lines. Mutual information sees any dependence — and costs an estimator.
Fig 2.16Why the choice of filter criterion decides which features survive. All three panels contain a strong, obviously usable relationship. Pearson correlation reports one of them and misses the other two entirely, so a correlation-based filter would discard the quadratic and circular features before the model ever saw them. Mutual information, I(X;Y) = H(Y) − H(Y|X), detects all three — but it has to be estimated from finite samples, and that estimate degrades quickly as dimensionality rises.
  • Wrapper methods search over subsets using the model itself as the scoring function. Recursive feature elimination fits the model, ranks features by an importance measure, removes the weakest, and repeats. Expensive but accounts for interactions and redundancy.
  • Embedded methods select during fitting. Lasso minimises ‖y − Xβ‖² + λ‖β‖₁; the L1 penalty's corner geometry drives coefficients exactly to zero, so the fit and the selection are one problem. Tree ensembles give impurity- or permutation-based importances — prefer permutation importance, since impurity importance is biased toward high-cardinality features.
The hygiene rule that decides whether a result is valid

Feature selection is part of the model, so it must happen inside each cross-validation fold, using only that fold's training data. Selecting features on the full dataset and then cross-validating the classifier leaks label information and produces optimistic estimates — routinely several accuracy points, and much more when the number of features exceeds the number of samples. The same applies to normalisation statistics, PCA bases, and class-balancing.

2.16Dimensionality reduction

PCA in one derivation

Centre the data, X̃ = X − μ. Seek the unit direction w maximising the projected variance wTΣw subject to wTw = 1, where Σ = X̃TX̃/(n−1). The Lagrangian gives Σw = λw: the optimal directions are the eigenvectors of the covariance matrix, and the variance captured by each is its eigenvalue. Keep the top d so that Σi≤dλi / Σiλi reaches your threshold. In practice compute it by SVD of X̃ rather than by forming Σ, which squares the condition number.

On device PCA is just a matrix multiply by a stored D×d matrix — cheap at inference, but the matrix itself occupies flash, which can exceed what the reduction saves. Always compare D·d stored weights against the parameters saved downstream.

-4 0 4 -4 0 4 PC1 PC2 260 samples, two correlated features PC1 PC2 0 % 50 % 100 % 96 % 4 % variance explained Keep PC1 → 1 number instead of 2, losing 4 % of the variance. PCA finds the eigenvectors of the covariance matrix; the eigenvalue is the variance along each. On device it is one stored matrix multiply — count its bytes before assuming it saves any.
Fig 2.17PCA in two dimensions, computed on 260 samples. The axes are the eigenvectors of the covariance matrix, drawn with lengths proportional to the square roots of their eigenvalues. Note the embedded warning: the projection matrix itself has to live in flash. For a 512-dimensional feature reduced to 32, that is 16 384 stored coefficients — often more than the classifier you were trying to shrink.

LDA is supervised: it maximises the Fisher criterion wTSBw / wTSWw, the ratio of between-class to within-class scatter. Its output is capped at C−1 dimensions for C classes — a hard structural limit worth remembering — and it assumes roughly Gaussian, equal-covariance classes.

t-SNE and UMAP are non-linear neighbour-embedding methods for visualisation. Both optimise a notion of local neighbourhood preservation; both distort global distances; t-SNE has no natural out-of-sample extension at all. They are excellent for inspecting whether your classes separate before you spend a week training, and they must never appear in a deployed pipeline. Beware of reading cluster sizes and inter-cluster distances in a t-SNE plot: neither is meaningful.

Random projection deserves more attention than it gets in embedded work. The Johnson–Lindenstrauss lemma guarantees that projecting onto O(log n / ε²) random directions preserves all pairwise distances to within 1±ε. With a sparse or ±1 projection matrix generated from a stored seed, the projection costs additions only and needs no stored matrix — an unusually good fit for a microcontroller.

PCA versus LDA, and t-SNE, on pictures

The two figures below make the comparison of this section concrete. In the first, the direction of greatest variance is exactly the wrong one to keep for classification; LDA, which uses the labels, finds the one that separates the classes. Both are a fixed matrix multiply at inference time, with parameters estimated on the training split only.

PCA direction LDA direction projection on the PCA direction — classes overlap projection on the LDA direction — classes separate PCA ignores labels; LDA uses them (supervised).
Fig 2.18The direction of greatest variance is not necessarily the direction that separates the classes. PCA (dashed) finds the long axis of the whole cloud and, projected onto it, the two classes are indistinguishable; Fisher's linear discriminant (copper) uses the labels to find the direction that maximises between-class over within-class scatter. Both are a single dot product per output dimension at inference time.

The second figure shows the same contrast between a linear projection and a non-linear embedding on real data. Use the t-SNE view to understand a dataset — whether classes form clusters, which examples are mislabelled, whether one subject's recordings sit apart from the rest — and never as a stage of the deployed pipeline.

PCA — linear, global, 2 of 64 dimensions 0 1 2 3 4 5 6 7 8 9 t-SNE — non-linear, local neighbourhoods 0 1 2 3 4 5 6 7 8 9 scikit-learn digits dataset (8×8 images = 64 features), 700 random samples; colour and numeral = true class. t-SNE axes mean nothing and between-cluster distances are not preserved: use it to look.
Fig 2.19Linear versus non-linear dimensionality reduction on real data. PCA keeps the two directions of largest variance and is a fixed matrix multiply — cheap enough to run on a microcontroller as a preprocessing step. t-SNE (and UMAP) preserve local neighbourhoods and reveal the cluster structure, but they have no cheap out-of-sample transform and distort global distances: they are tools for understanding a dataset on a workstation, not for deployment.
If you remember one thing from Session 2

Decide what your model must be able to distinguish, then work backwards: the bandwidth sets the sampling rate, the dynamic range sets the bits, the event duration sets the window, and the cheapest representation that keeps the distinguishing information is the right one. Anything lost before the model is lost for good — and anything kept that the model does not need is paid for in every inference.

Before Session 3
  • Y. Zhang, N. Suda, L. Lai, V. Chandra, "Hello Edge: Keyword Spotting on Microcontrollers," 2017. arXiv:1711.07128.A complete embedded pipeline — MFCC front end, several model families, memory and operation budgets — in twelve pages. Session 3 uses its numbers.
  • If you have no ML background: work through the primer sections at the start of Session 3 before the lecture.
Discussion questions
  1. A colleague proposes to sample an accelerometer at 25 Hz "because human motion is below 10 Hz". What do you need to know about the sensor's internal filter before agreeing?
  2. Name a task for which discarding phase (taking the magnitude spectrogram) destroys the information the model needs.
  3. When is a learned front end (a 1-D convolution on raw samples) likely to be cheaper overall than an FFT plus mel filterbank?
  4. Your test accuracy drops from 97 % to 81 % when you change from a random split to a subject-wise split. What have you learned, and what do you report?
Exercises
  1. A 12-bit ADC digitises a sensor whose useful signal spans 20 % of the full-scale range. What is the effective SQNR for that signal, and how many bits are effectively "used"?
  2. Compute the number of frames, the frequency resolution and the output size of a 512-point STFT with hop 160 on one second of 16 kHz audio. How do they change with a 1024-point frame?
  3. An IMU samples at 50 Hz. You need to recognise gestures lasting 0.6–1.5 s with a decision latency below 1 s. Propose window length, hop and three features, and estimate the compute per second.
  4. Derive the cost in MACs of computing 40 log-mel energies from a 512-point power spectrum when the filterbank is stored sparsely.
Further reading
  • A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, 3rd ed., Pearson, 2010.The reference for sampling, filtering and the DFT.
  • P. Warden, "Speech Commands," 2018. arXiv:1804.03209.Dataset design for embedded audio.
  • D. S. Park et al., "SpecAugment," Interspeech 2019. arXiv:1904.08779.The standard augmentation for spectrogram inputs.
  • Q. Kong et al., "PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition," IEEE/ACM TASLP 2020. arXiv:1912.10211.How far learned audio representations go when compute is plentiful — the opposite end of the spectrum from this course.