On-Chip Memory Optimization for Deep Learning Inference on Resource-Deficient Devices

Deepak Kumar, Udit Satija · IEEE transactions on circuits and systems for artificial intelligence. · 2024

Advances in deep learning (DL) have enabled the integration of intelligence into low-end Internet-of-Things (IoT) devices. However, traditional DL inference on resource-constrained microcontrollers (MCUs) faces significant memory and computational challenges. This article presents a novel compression and reconstruction method for on-chip memory optimization enabling the DL model on such MCUs. We exploit a low-rank approximation technique to decompose the deep neural network (DNN) weight matrix into lower-rank sub-matrices that occupy less Flash memory than the uncompressed matrix. Further, we propose a novel layer-wise kernel reconstruction (LaWKeR) technique for static random access memory (SRAM) and FLOPS reduction. The LaWKeR reconstructs only the active kernel directly from compressed weights stored in Flash, unlike existing methods which reconstruct the whole compressed weight matrix during run-time. We demonstrate the efficacy of the proposed method on a NUCLEO board using the STM32F401 microcontroller. The LaWKeR method reduces the model size from 73KB to 20.4KB of Flash and RAM usage of only 21.34KB, using only 4.20% and 21.28% of the total embedded memory, respectively. It requires less memory than popular frameworks like tinyML and neural architecture search and can be combined with other techniques for even smaller models.

Read the paper · More papers on PaperTik