Optimizing Mixed-Precision DNN Scheduling on Heterogeneous SoCs for Enhanced Robustness and Efficiency
Yulong Song, Qiang Tao, Jinlun Ji, Congyi Sun, Yuxiang Fu, Li Li · 2025
Deep Neural Network (DNN) inference workloads, such as object recognition in autonomous systems, place significant demands on performance and energy efficiency. Additionally, real-world DNN inference workloads are usually influenced by external factors, such as bad weather, which can lead to a degradation in accuracy. To address performance and power issues, techniques such as quantization and model pruning have been introduced, but these can worsen model robustness. This study explores mixed-precision DNN inference on a heterogeneous System-on-Chip (SoC) to balance performance and power consumption while maintaining model robustness. We allocate each DNN layer to either a performance- or energy-efficient accelerator, utilizing different computational precisions. In our work, a performance cost model is formutaled first to predict the total execution time. Then, we determine the schedules using reinforcement learning, with the cost model providing feedbacks. Our approach is evaluated on the NVIDIA Xavier SoC with commonly used DNN models, achieving 97.76% of full-precision accuracy, while reducing inference latency and energy consumption by up to 16% and 33%, respectively, compared to the state-of-the-art.