EFFICIENT DEPLOYMENT OF MACHINE LEARNING MODELS ON MICROCONTROLLERS: A COMPARATIVE STUDY OF QUANTIZATION AND PRUNING STRATEGIES.

Rafael Bessa Loureiro, Paulo Henrique Miranda Sá, Fernanda Vitoria Nascimento Lisboa, Rodrigo Matos Peixoto, Lian Filipe Santana Nascimento, Yasmin da Silva Bonfim, Gustavo Oliveira Ramos Cruz, T. Ramos, Carlos Henrique Racobaldo Luz Montes, Tiago Palma Pagano, Oberdan Rocha Pinheiro, Rafael V. Borges · 2023

With the advancement and growth of Internet of Things tools, the necessity for more complex and intelligent systems increases, presenting many challenges due to device limitations in memory, computation, and energy consumption. The objective of this study is to do a literature review of optimization techniques with quantization and pruning in machine learning models for deployment on microcontrollers. The methodology consists of searching the literature, the tools, the models, and the techniques. Results show that pruning has better accuracy, while quantization has better inference time and size reduction, with the best results reducing 90% with 3.5x faster inference. Careful consideration of trade-offs between inference time, model size, and accuracy is crucial for deployment on edge devices.

Read the paper · More papers on PaperTik