Deep Learning Compiler Optimization Based on Constraint Programming

Lucheng Xie · 2024

In contemporary computer vision research, optimizing the inference performance of Convolutional Neural Networks (CNNs) on Graphics Processing Units (GPUs) represents a substantial challenge. This challenge primarily emanates from the need to fully utilize hardware resources and efficiently schedule complex computational tasks. Current deep learning compilers, including RAMMER, strive to maximize hardware resource utilization by leveraging parallelism through cooperative scheduling across operators. Yet, this conventional scheduling approach often falls short in fully exploiting the capabilities of GPU parallel processing. This study introduces a novel scheduling methodology utilizing constraint programming technology to enhance the CNN inference performance on GPUs. Our approach integrates multiple operators into a single composite operator, efficiently exploiting both inter-operator and intra-operator parallelism. This strategy enhances hardware resource utilization, thereby optimizing overall inference performance. In our experimental evaluation, the proposed constraint programming-based scheduling method was benchmarked against the traditional RAMMER scheduling strategy. The findings reveal that our method consistently surpasses RAMMER in inference performance by approximately 1.1 to 1.2 times, across a spectrum of CNN models and diverse computational environments.

Read the paper · More papers on PaperTik