Bibliography
Foundations, surveys and systems
- V. Sze, Y.-H. Chen, T.-J. Yang, J. Emer. "Efficient Processing of Deep Neural Networks: A Tutorial and Survey." Proc. IEEE, 2017.arXiv:1703.09039
- M. Horowitz. "1.1 Computing's Energy Problem (and what we can do about it)." ISSCC, 2014.Source of the energy-per-operation table.
- S. Williams, A. Waterman, D. Patterson. "Roofline: An Insightful Visual Performance Model for Multicore Architectures." CACM 52(4), 2009.
- K. Goto, R. van de Geijn. "Anatomy of High-Performance Matrix Multiplication." ACM TOMS 34(3), 2008.
- P. Warden, D. Situnayake. TinyML. O'Reilly, 2019.
- V. Janapa Reddi et al. Machine Learning Systems (open textbook), mlsysbook.ai.
- S. Somvanshi et al. "From Tiny Machine Learning to Tiny Deep Learning: A Survey." 2025.arXiv:2506.18927
Signals, features and audio
- P. Warden. "Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition." 2018.arXiv:1804.03209
- Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, M. D. Plumbley. "PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition." 2020.arXiv:1912.10211
- Y. Zhang, N. Suda, L. Lai, V. Chandra. "Hello Edge: Keyword Spotting on Microcontrollers." 2017.arXiv:1711.07128
- Y. Gong, Y.-A. Chung, J. Glass. "AST: Audio Spectrogram Transformer." Interspeech, 2021.arXiv:2104.01778
- B. Kim, S. Chang, J. Lee, D. Sung. "Broadcasted Residual Learning for Efficient Keyword Spotting." Interspeech, 2021.arXiv:2106.04140
- D. S. Park et al. "SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition." Interspeech, 2019.arXiv:1904.08779
- A. Chowdhery, P. Warden, J. Shlens, A. Howard, R. Rhodes. "Visual Wake Words Dataset." 2019.arXiv:1906.05721
- L. McInnes, J. Healy, J. Melville. "UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction." 2018.arXiv:1802.03426
- H. Peng, F. Long, C. Ding. "Feature Selection Based on Mutual Information: Criteria of Max-Dependency, Max-Relevance, and Min-Redundancy." IEEE TPAMI 27(8), 2005.
Classical models
- A. Rahimi, B. Recht. "Random Features for Large-Scale Kernel Machines." NIPS, 2007.
- T. Hastie, R. Tibshirani, J. Friedman. The Elements of Statistical Learning, 2nd ed. Springer, 2009.
- T. Chen, C. Guestrin. "XGBoost: A Scalable Tree Boosting System." KDD, 2016.
Architectures and search
- F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, K. Keutzer. "SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size." 2016.arXiv:1602.07360
- A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, H. Adam. "MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications." 2017.arXiv:1704.04861
- M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen. "MobileNetV2: Inverted Residuals and Linear Bottlenecks." CVPR, 2018.arXiv:1801.04381
- A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan, Q. V. Le, H. Adam. "Searching for MobileNetV3." ICCV, 2019.arXiv:1905.02244
- X. Zhang, X. Zhou, M. Lin, J. Sun. "ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices." CVPR, 2018.arXiv:1707.01083
- N. Ma, X. Zhang, H.-T. Zheng, J. Sun. "ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design." ECCV, 2018.arXiv:1807.11164
- M. Tan, Q. V. Le. "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks." ICML, 2019.arXiv:1905.11946
- J. Hu, L. Shen, S. Albanie, G. Sun, E. Wu. "Squeeze-and-Excitation Networks." CVPR, 2018.arXiv:1709.01507
- S. Mehta, M. Rastegari. "MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer." ICLR, 2022.arXiv:2110.02178
- M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, F. S. Khan. "EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications." 2022.arXiv:2206.10589
- T. Elsken, J. H. Metzen, F. Hutter. "Neural Architecture Search: A Survey." JMLR 20, 2019.arXiv:1808.05377
- M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, Q. V. Le. "MnasNet: Platform-Aware Neural Architecture Search for Mobile." CVPR, 2019.arXiv:1807.11626
- H. Cai, L. Zhu, S. Han. "ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware." ICLR, 2019.arXiv:1812.00332
- H. Cai, C. Gan, T. Wang, Z. Zhang, S. Han. "Once-for-All: Train One Network and Specialize it for Efficient Deployment." ICLR, 2020.arXiv:1908.09791
- C. Banbury, C. Zhou, I. Fedorov, R. Matas Navarro, U. Thakker, D. Gope, V. Janapa Reddi, M. Mattina, P. N. Whatmough. "MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers." MLSys, 2021.arXiv:2010.11267
- S. Bai, J. Z. Kolter, V. Koltun. "An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling." 2018.arXiv:1803.01271
- A. Gu, T. Dao. "Mamba: Linear-Time Sequence Modeling with Selective State Spaces." 2023.arXiv:2312.00752
Quantization
- B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, D. Kalenichenko. "Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference." CVPR, 2018.arXiv:1712.05877
- M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. van Baalen, T. Blankevoort. "A White Paper on Neural Network Quantization." 2021.arXiv:2106.08295
- R. Krishnamoorthi. "Quantizing deep convolutional networks for efficient inference: A whitepaper." 2018.arXiv:1806.08342
- A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, K. Keutzer. "A Survey of Quantization Methods for Efficient Neural Network Inference." 2021.arXiv:2103.13630
- M. Nagel, M. van Baalen, T. Blankevoort, M. Welling. "Data-Free Quantization Through Weight Equalization and Bias Correction." ICCV, 2019.arXiv:1906.04721
- M. Nagel, R. A. Amjad, M. van Baalen, C. Louizos, T. Blankevoort. "Up or Down? Adaptive Rounding for Post-Training Quantization." ICML, 2020.arXiv:2004.10568
- S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, D. S. Modha. "Learned Step Size Quantization." ICLR, 2020.arXiv:1902.08153
- Y. Bengio, N. Léonard, A. Courville. "Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation." 2013.arXiv:1308.3432 — the straight-through estimator.
- M. Courbariaux, Y. Bengio, J.-P. David. "BinaryConnect: Training Deep Neural Networks with binary weights during propagations." NIPS, 2015.arXiv:1511.00363
- M. Rastegari, V. Ordonez, J. Redmon, A. Farhadi. "XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks." ECCV, 2016.arXiv:1603.05279
- Z. Dong, Z. Yao, A. Gholami, M. Mahoney, K. Keutzer. "HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision." ICCV, 2019.arXiv:1905.03696
Pruning, sparsity and distillation
- Y. LeCun, J. S. Denker, S. A. Solla. "Optimal Brain Damage." NIPS, 1990.
- B. Hassibi, D. G. Stork. "Second Order Derivatives for Network Pruning: Optimal Brain Surgeon." NIPS, 1993.
- S. Han, J. Pool, J. Tran, W. J. Dally. "Learning both Weights and Connections for Efficient Neural Networks." NIPS, 2015.arXiv:1506.02626
- S. Han, H. Mao, W. J. Dally. "Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding." ICLR, 2016.arXiv:1510.00149
- J. Frankle, M. Carbin. "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks." ICLR, 2019.arXiv:1803.03635
- Z. Liu, M. Sun, T. Zhou, G. Huang, T. Darrell. "Rethinking the Value of Network Pruning." ICLR, 2019.arXiv:1810.05270
- H. Li, A. Kadav, I. Durdanovic, H. Samet, H. P. Graf. "Pruning Filters for Efficient ConvNets." ICLR, 2017.arXiv:1608.08710
- Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, C. Zhang. "Learning Efficient Convolutional Networks through Network Slimming." ICCV, 2017.arXiv:1708.06519
- M. Zhu, S. Gupta. "To prune, or not to prune: exploring the efficacy of pruning for model compression." 2017.arXiv:1710.01878
- Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, S. Han. "AMC: AutoML for Model Compression and Acceleration on Mobile Devices." ECCV, 2018.arXiv:1802.03494
- T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, A. Peste. "Sparsity in Deep Learning." JMLR, 2021.arXiv:2102.00554
- G. Hinton, O. Vinyals, J. Dean. "Distilling the Knowledge in a Neural Network." NIPS Deep Learning Workshop, 2014.arXiv:1503.02531
- A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, Y. Bengio. "FitNets: Hints for Thin Deep Nets." ICLR, 2015.arXiv:1412.6550
- J. Gou, B. Yu, S. J. Maybank, D. Tao. "Knowledge Distillation: A Survey." IJCV, 2021.arXiv:2006.05525
- H. Zhang, M. Cisse, Y. N. Dauphin, D. Lopez-Paz. "mixup: Beyond Empirical Risk Minimization." ICLR, 2018.arXiv:1710.09412
Hardware and accelerators
- N. P. Jouppi et al. "In-Datacenter Performance Analysis of a Tensor Processing Unit." ISCA, 2017.arXiv:1704.04760
- Y.-H. Chen, T. Krishna, J. S. Emer, V. Sze. "Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks." IEEE JSSC 52(1), 2017.
- S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, W. J. Dally. "EIE: Efficient Inference Engine on Compressed Deep Neural Network." ISCA, 2016.arXiv:1602.01528
- L. Lai, N. Suda, V. Chandra. "CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs." 2018.arXiv:1801.06601
- A. Lavin, S. Gray. "Fast Algorithms for Convolutional Neural Networks." CVPR, 2016.
- M. Le Gallo et al. "A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference." Nature Electronics 6, 2023.
TinyML systems, benchmarking and deployment
- C. R. Banbury, V. Janapa Reddi, M. Lam, et al. "Benchmarking TinyML Systems: Challenges and Direction." 2020.arXiv:2003.04821
- C. Banbury, V. Janapa Reddi, P. Torelli, J. Holleman, N. Jeffries, C. Kiraly, et al. "MLPerf Tiny Benchmark." 2021.arXiv:2106.07597
- J. Lin, W.-M. Chen, Y. Lin, J. Cohn, C. Gan, S. Han. "MCUNet: Tiny Deep Learning on IoT Devices." NeurIPS, 2020.arXiv:2007.10319
- J. Lin, W.-M. Chen, H. Cai, C. Gan, S. Han. "MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning." NeurIPS, 2021.arXiv:2110.15352
- R. David, J. Duke, A. Jain, V. Janapa Reddi, N. Jeffries, J. Li, N. Kreeger, I. Nappier, M. Natraj, S. Regev, R. Rhodes, T. Wang, P. Warden. "TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems." MLSys, 2021.arXiv:2010.08678
- J. Lin, L. Zhu, W.-M. Chen, W.-C. Wang, C. Gan, S. Han. "On-Device Training Under 256KB Memory." NeurIPS, 2022.arXiv:2206.15472
- T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, M. Cowan, H. Shen, L. Wang, Y. Hu, L. Ceze, C. Guestrin, A. Krishnamurthy. "TVM: An Automated End-to-End Optimizing Compiler for Deep Learning." OSDI, 2018.arXiv:1802.04799
- H. B. McMahan, E. Moore, D. Ramage, S. Hampson, B. Agüera y Arcas. "Communication-Efficient Learning of Deep Networks from Decentralized Data." AISTATS, 2017.arXiv:1602.05629
- E. Frantar, S. Ashkboos, T. Hoefler, D. Alistarh. "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers." ICLR, 2023.arXiv:2210.17323
- J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, S. Han. "AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration." MLSys, 2024.arXiv:2306.00978
Comparable courses consulted in designing this syllabus
- MIT 6.5940, TinyML and Efficient Deep Learning Computing (S. Han), hanlab.mit.edu.Source of the depth of the compression treatment in Session 3.
- Harvard CS249r, Tiny Machine Learning (V. Janapa Reddi), and the associated Machine Learning Systems textbook.Source of the systems-and-benchmarking emphasis in Session 4.
- ETH Zürich 227-0155-00L, Machine Learning on Microcontrollers.Source of the signal-chain depth in Session 2.
Textbooks and reference works
- A. V. Oppenheim, R. W. Schafer. Discrete-Time Signal Processing, 3rd ed. Pearson, 2010.
- J. L. Hennessy, D. A. Patterson. Computer Architecture: A Quantitative Approach, 6th ed. Morgan Kaufmann, 2017.
- V. Sze, Y.-H. Chen, T.-J. Yang, J. Emer. Efficient Processing of Deep Neural Networks. Morgan & Claypool, 2020.
- I. Goodfellow, Y. Bengio, A. Courville. Deep Learning. MIT Press, 2016 (free online at deeplearningbook.org).For participants who want a full treatment of the primer material of Session 3.
- C. M. Bishop, H. Bishop. Deep Learning: Foundations and Concepts. Springer, 2024.