Optimizing Machine Learning Models Using Tensor Virtual Machine for Embedded CPUs

Ashish Tiwari, Christophe Fava-Rivi, Sahil Munaf Bandar, Sahil Salim Makandar · 2024

Optimization of deep learning models for embedded CPUs presents numerous challenges stemming from limited computational resources, memory constraints, thread synchronization overhead, and the imperative of efficient Single Instruction Multiple Data (SIMD) instruction utilization. Tailored optimizations are essential to achieve peak performance on such hardware platforms. This research focuses on optimizing pre-trained deep learning models for embedded CPUs such as Cortex A53 and Cortex A72 in Rockchip RK3399ProD and NXP i.MX8MPlus. The study employs YOLOv3 and Tiny YOLOv3 models pretrained on the COCO dataset for object detection tasks to demonstrate the efficiency of the Tensor Virtual Machine compiler toolchain. TVM, an open-source compiler stack for deep learning, facilitates the acceleration of machine learning models across diverse hardware platforms by optimizing model performance. Experimental results indicate that integrating TVM into the multi-step conversion process yielded a significant performance enhancement, achieving approximately ~2 to 3 times higher frames per second (fps) on an embedded NXP i.MX8MPlus.

Read the paper · More papers on PaperTik