Inference at the Edge for Complex Deep Learning Applications with Multiple Models and Accelerators
K P Ashwanth, Dharun Narayanan L K, Dev Divyendh D, Sreehari Krishna S, Vijaya Kumar Sundar, Priyanka D. Kumar · 2023
In this paper, we demonstrate the performance benefits of offloading deep learning workloads to specialized hardware accelerators using an experimental setup consisting of a Raspberry Pi 4 Model B board and two Intel Movidius Neural Compute Sticks. Our study shows that the performance of complex deep learning applications, which require a cluster of models to achieve their goals, can be improved by assigning suitable models to accelerators based on their computational requirements. We propose a two-step algorithm that analyzes the computational requirements of each model by taking their .xml and .bin files as input along with the number of accelerators interfaced with the edge device. In the first step, the algorithm calculates the computational complexity score of each model based on its number of tags and edge tags. In the second step, the algorithm allocates the models to the available accelerators based on their current load and computational complexity score. Our findings demonstrate the effectiveness of Intel Movidius Neural Compute Sticks in increasing frames per second while reducing latency for deep learning functions implemented on peripheral devices like the Raspberry Pi.