ConCo: Optimizing Compilation of Concurrent Tensor Programs on Shared GPU

Jiamin Lu, Jingwei Sun, Yunlong Xu, Peng Sun, Guangzhong Sun · 2025

Serving multiple inference tasks of deep neural networks (DNNs) concurrently on a shared GPU is an established method for maximizing hardware resource.Although DNN compilers effectively generate optimal kernel code for individual DNN inferences, they fall short in optimizing for concurrent tasks.This paper presents ConCo, a concurrencyaware compilation scheme designed to optimize the execution of concurrent DNN inference tasks on a shared GPU.ConCo dynamically generates multiple code variants, each tailored to different GPU resource constraints, and efficiently selects optimal variants at runtime according to concurrent workload characteristics.To mitigate the substantial overhead associated with multi-variant compilation, ConCo employs an optimal-code-sharing strategy, significantly accelerating compilation by leveraging commonalities across resource configurations.Evaluations demonstrate that ConCo improves inference throughput by up to 1.2× and reduces job completion time by up to 69.85% compared to existing solutions.

Read the paper · More papers on PaperTik