Extensions for compiler infrastructure for neural networks to support automated generation of high-performance tensor programs for DSPs
Rui Huang, Zhe Li, Yuan Ji, Jing Yang, Yaohua Wang · 2024
To widely apply deep learning technology, it is necessary to deploy neural network models on different hardware platforms. Multi-core Digital Signal Processors (DSPs), as one of these platforms, exhibit the potential to accelerate deep learning in specific applications and scenarios due to their energy efficiency advantages. CINN, as an AI compiler, is been proposed for deploying DNN models on diverse hardware devices. However, due to differences in hardware architectures, there is a risk of poor utilization of memory and computing resources, leading to increased execution time. In this work, we propose several extensions to achieve optimization specific to DSPs, and allow CINN to automatically generate high-performance tensor programs for our DSPs through a search approach. Experimental result shows that there is a significant improvement in execution performance for the automatically generated tensor operators containing reduction operation compared to manual schedule.