EmbML · S2 · Data & features1 / 48

Embedded Machine Learning · PhD course

Session 02

Data acquisition, preprocessing and feature engineering

What has already been lost before the model sees the data — and the cheapest representation that keeps what it needs.

90 minutesSession 2 of 5
EmbML · S2 · Data & features2 / 48

The principle

The model can only ever recover what the chain preserved.

Two steps in this session are irreversible — aliasing and clipping. Everything else can be revisited later.

EmbML · S2 · Data & features3 / 48

Overview

The chain, and what each stage destroys

The analog/digital boundary sits between the anti-alias filter and the sampler — and that placement is the whole point.

EmbML · S2 · Data & features4 / 48

The sensors you will meet

Rate, resolution, interface, power

SensorTypical rateBitsInterfaceActive power
MEMS microphone16–48 kHz16–24PDM / I²S≈ 0.3–1 mW
Accelerometer10 Hz – 6.4 kHz12–16I²C / SPI≈ 5–200 µW
Gyroscope25 Hz – 6.4 kHz16I²C / SPI≈ 1–3 mW
PPG (optical HR)25–500 Hz16–20I²C≈ 0.1–2 mW
ECG front end125–1000 Hz12–24SPI / ADC≈ 10–300 µW
Low-res camera1–30 fps8 / pixelDCMI / MIPI≈ 1–100 mW

The gyroscope costs 10–100× the accelerometer. And the sensor's configuration is part of the model's input spec.

EmbML · S2 · Data & features5 / 48

01

Signals in time and frequency

A primer — the language every later step is described in.

EmbML · S2 · Data & features6 / 48

Fourier

Any periodic signal is a sum of sinusoids

EmbML · S2 · Data & features7 / 48

Primer · the DFT

Five lines on the discrete Fourier transform

X[k] = Σₙ x[n] · e^(−j2πkn/N), k = 0 … N−1
bin frequency f_k = k · f_s / N resolution Δf = f_s / N
real signal: N/2+1 independent bins, highest frequency f_s / 2
FFT cost ≈ N·log₂N (N = 512: ~4.6 k vs N² ≈ 262 k)

Each X[k] measures how strongly the block correlates with a sinusoid at f_k. A longer block resolves finer frequency differences — at the price of time resolution.

EmbML · S2 · Data & features8 / 48

Primer · convolution

Why frequency is the natural language of the chain

  • A linear filter = convolution with its impulse response: a sliding weighted sum.
  • Convolution in time = multiplication in frequency: filtering scales each frequency.
  • And the reverse: sampling multiplies by an impulse train → the spectrum is copied every f_s.
  • Same word in Session 3: a CNN layer is a sliding weighted sum with learned weights.
EmbML · S2 · Data & features9 / 48

01

Sampling and aliasing

Stated precisely, then demonstrated numerically.

EmbML · S2 · Data & features10 / 48

The theorem, and what it really says

Sampling replicates the spectrum

Reconstructable ⟺ f_s > 2B (B = highest frequency present)
f_alias = | f − round(f / f_s) · f_s |

Sampling in time multiplies by an impulse train, which convolves the spectrum with an impulse train. The spectrum is periodically replicated. If f_s < 2B the replicas overlap and ADD — and an addition whose terms were never observed apart cannot be undone.

EmbML · S2 · Data & features11 / 48

Mechanism

Aliasing is spectral folding

Replicas overlap → the observed spectrum is their sum.

EmbML · S2 · Data & features12 / 48

Computed, not asserted

Same samples, two different signals

A 7 Hz tone sampled at 10 Hz produces exactly the sample set a 3 Hz tone would produce.

EmbML · S2 · Data & features13 / 48

02

Quantization noise

Where the second irreversible loss happens — and how much of it you are allowed.

EmbML · S2 · Data & features14 / 48

Derivation

From a uniform quantizer to 6.02 N + 1.76 dB

q = V_FS / 2^N step size
σ²_e = (1/q) ∫ e² de over [−q/2, q/2] = q²/12
SQNR = 10 log₁₀ ( (V_FS²/8) / (q²/12) ) ≈ 6.02 N + 1.76 [dB]

Assumption: the rounding error is uniform on [−q/2, q/2] and independent of the signal. Good when the signal is busy relative to q — and false when it is not.

EmbML · S2 · Data & features15 / 48

What a quantizer really does

The error is a sawtooth, not noise

Every flat tread is a set of real inputs the converter cannot tell apart.

EmbML · S2 · Data & features16 / 48

Design against this, not the label

Bits are worth 6 dB — in both directions

ENOB = (SINAD − 1.76) / 6.02. A “16-bit” ADC on a noisy board often measures 11.

EmbML · S2 · Data & features17 / 48

Foundations

Three converter architectures, three currencies

ArchitectureMechanismTypicalImplication for you
SARbinary search, one comparison per bit8–18 bit
kSPS–MSPS
The MCU default. Low latency, easy to multiplex across channels.
Sigma–delta1-bit modulator at 64–256× oversampling + digital decimation16–24 bit
kSPS
Audio and precision sensing. Adds group delay — matters for event timestamps and sensor fusion.
Flash / pipeline2ᴺ−1 parallel comparators, or a pipeline of stages6–14 bit
MSPS–GSPS
Radar, ultrasound, RF. You will be bandwidth-limited downstream long before the converter is the problem.

Each buys resolution with a different currency: comparisons, time, or silicon area.

EmbML · S2 · Data & features18 / 48

03

Noise and filtering

What actually limits your SNR — and it is rarely the quantizer.

EmbML · S2 · Data & features19 / 48

Choosing a structure

FIR versus IIR on an embedded target

PropertyFIRIIR
Cost per output sampleM MACs (M taps)≈ 2 × order (biquads)
State memoryM words2 words per biquad
Phaseexactly linear if symmetricnon-linear; group delay varies
Stabilityalways stablepoles must stay inside the unit circle
Fixed-point failureaccumulator overflowcoefficient quantization moves poles; limit cycles
Use fordecimation, matched filtering, anything phase-sensitivecheap band-limiting, DC blocking, envelopes

For a given magnitude spec an IIR is typically an order of magnitude cheaper. You pay in phase linearity and numerical fragility.

EmbML · S2 · Data & features20 / 48

Two cheap filters

Moving average versus median

EmbML · S2 · Data & features21 / 48

Sampling rate

Halving f_s halves almost everything downstream

  • ADC energy per conversion × rate
  • Filter MACs per sample × rate
  • Frames per second × feature cost
  • Inferences per second × model cost
  • But: only free of accuracy loss if the discarded band carries no class-discriminative information — test it, do not assume it.

Keyword spotting: 16 kHz conventional, 8 kHz often enough. Bearing faults: the diagnostic energy may sit above 10 kHz and the same reasoning gives the opposite answer.

EmbML · S2 · Data & features22 / 48

Break

Five minutes.

Next: datasets and preprocessing, then features.

EmbML · S2 · Data & features23 / 48

02

Datasets, labels and preprocessing

The most expensive and least reusable part of most projects.

EmbML · S2 · Data & features24 / 48

Building a dataset

Coverage, labels, rarity

  • Coverage beats volume: users, mountings, devices, environments.
  • Label granularity is a design decision: per recording, per window, exact boundaries.
  • Rare events need targeted collection and a diverse background class.
  • Preprocessing — sync, resample, gaps, calibration — must exist identically in firmware.
  • Augment: noise, gain, shifts, rotations, SpecAugment. Free at inference.
EmbML · S2 · Data & features25 / 48

Normalisation

Scale features before most models see them

EmbML · S2 · Data & features26 / 48

The commonest invalid result

Split by subject, device or recording — never by window.

Overlapping windows from one recording are near-copies. A random split measures memorisation.

EmbML · S2 · Data & features27 / 48

Bridge to Session 2

Framing: the model consumes frames, not streams

  • Frame length L, hop H → overlap 1 − H/L, frame rate f_s/H
  • Frame rate sets the inference rate, and therefore the energy
  • Frame length sets the lowest resolvable frequency, f_s/L
  • Hop sets the temporal precision of an event onset
  • Overlap is not free — 50 % overlap doubles the number of inferences

On an energy-limited device, overlap is a parameter to be justified, never a default.

EmbML · S2 · Data & features28 / 48

Segmentation

From a stream to examples

EmbML · S2 · Data & features29 / 48

Method

Three questions to ask of any representation

  • What does it discard — and is the discarded part class-discriminative?
  • What invariances does it build in? A magnitude spectrogram discards phase, buying shift invariance. If your task needs inter-channel phase, you just deleted the signal.
  • What does it cost, compared with the model it feeds?

If the front end costs 40 % of the budget, a learned front end on raw samples may be cheaper overall.

EmbML · S2 · Data & features30 / 48

Time-domain features

Energy and zero-crossing rate: a voice-activity detector for free

EmbML · S2 · Data & features31 / 48

01

The short-time Fourier transform

Windows, leakage, and a bound you cannot tune your way around.

EmbML · S2 · Data & features32 / 48

Definition

Windowed transforms of overlapping segments

X[m, k] = Σ_n x[n] · w[n − mH] · exp(−j2πkn/L)
spectrogram = |X[m,k]|²

The Fourier transform assumes stationarity over its whole support. Speech, vibration and gesture signals are not stationary — so we apply it to short segments in which stationarity is approximately true.

EmbML · S2 · Data & features33 / 48

Measured, not asserted

Narrow main lobe or low sidelobes — not both

Rectangular −13 dB · Hann −31 · Hamming −43 · Blackman–Harris −92.

EmbML · S2 · Data & features34 / 48

The bound

Every tile has the same area

  • Δf ≈ f_s / L    Δt = L / f_s    Δt·Δf ≈ 1
  • You cannot resolve a 20 ms transient and a 10 Hz frequency difference in the same analysis.
  • Overlap increases the frame rate, not the resolution.
  • The smearing is set by L. Students conflate these constantly.
EmbML · S2 · Data & features35 / 48

02

From spectrogram to mel and MFCC

Four steps, each of which had a reason — and one of which no longer applies.

EmbML · S2 · Data & features36 / 48

Perceptual warping

Resolution spent where hearing has it

m(f) = 2595 · log₁₀(1 + f/700). Six filters below 1 kHz; five across the whole 4–8 kHz span.

EmbML · S2 · Data & features37 / 48

The pipeline, and the branch that matters

Where to stop for a CNN

The DCT existed to decorrelate for diagonal-covariance GMMs. A CNN has no such requirement — and the DCT destroys the local structure convolutions exploit.

EmbML · S2 · Data & features38 / 48

The same sound, two representations

6.4× fewer numbers, nearly all the structure

Voiced harmonics, a fricative burst, a click — all still legible after mel compression.

EmbML · S2 · Data & features39 / 48

Price the front end

Front end against model, order of magnitude

Front endCost / frameOutput dimNote
Time-domain statistics, 3-axis IMU, 128 samples~10³ ops~30Runs in the sensor ISR
512-point real FFT~10⁴ MACs257CMSIS-DSP radix-2/4
40-band log-mel from that FFT~1.5 × 10⁴40Filterbank is sparse
13 MFCC + Δ + ΔΔ~1.6 × 10⁴39Adds a small DCT
Small DS-CNN over a 49×10 map~2.7 × 10⁶—Per inference ≈ 4× the front end of a whole second (49 frames ≈ 7 × 10⁵)

The comparison flips for very small models and high sample rates. Measure, do not assume.

EmbML · S2 · Data & features40 / 48

03

Selection and reduction

Which features survive — and the hygiene rule that decides whether a result is valid.

EmbML · S2 · Data & features41 / 48

Three families

Filter, wrapper, embedded

Filter — score each feature alone

  • Pearson ρ: linear dependence only
  • Mutual information: any dependence, but must be estimated
  • Cheap; blind to redundancy
  • Two perfectly correlated informative features both score high, both get kept

Wrapper & embedded

  • RFE: fit, rank, drop the weakest, repeat
  • Expensive, but accounts for interaction and redundancy
  • Lasso: L1 corner geometry drives coefficients to exactly zero
  • Tree importances — prefer permutation over impurity, which is biased toward high-cardinality features
EmbML · S2 · Data & features42 / 48

The hygiene rule

Feature selection is part of the model, so it happens inside each cross-validation fold.

Selecting on the full dataset and then cross-validating the classifier leaks label information — routinely several accuracy points, and far more when features outnumber samples. The same applies to normalisation statistics, PCA bases and class balancing.

EmbML · S2 · Data & features43 / 48

Dimensionality reduction

PCA, and the cost nobody counts

  • Maximise wᵀΣw subject to ‖w‖=1 → Σw = λw: the eigenvectors of the covariance matrix.
  • Compute it by SVD of the centred data, not by forming Σ — which squares the condition number.
  • On device it is one stored D×d matrix multiply.
  • Count its bytes. 512 → 32 is 16 384 stored coefficients — often more than the classifier you were shrinking.
EmbML · S2 · Data & features44 / 48

Supervised projection

Greatest variance is not greatest separation

EmbML · S2 · Data & features45 / 48

Looking at data

t-SNE is for looking, not for deploying

EmbML · S2 · Data & features46 / 48

Before Session 3

Reading

Y. Zhang, N. Suda, L. Lai, V. Chandra, “Hello Edge: Keyword Spotting on Microcontrollers,” 2017. arXiv:1711.07128.

Why: A complete embedded pipeline — MFCC front end, several model families, memory and operation budgets — in twelve pages. Session 3 uses its numbers.

No ML background? Work through Part A of the Session 3 web chapter before the lecture.

EmbML · S2 · Data & features47 / 48

If you remember one thing

Decide what the model must distinguish, then work backwards: bandwidth sets the sampling rate, dynamic range sets the bits, event duration sets the window, and the cheapest representation that keeps the distinguishing information is the right one.

EmbML · S2 · Data & features48 / 48

Before you go

Think about these

  1. A colleague samples an accelerometer at 25 Hz “because human motion is below 10 Hz”. What must you know about the sensor's internal filter before agreeing?
  2. Name a task for which discarding phase destroys the information the model needs.
  3. Your accuracy falls from 97 % to 81 % with a subject-wise split. What do you report?
S2 · Data & features