Reducing the End-to-End Latency of DNN-Based Recommendation Systems in GPU Pools

Guangqiang Luan, Pu Pang, Quan Chen, Chen Chen, Guoyao Xu, Chi Zhang, Yanyi Zi, Yinghao Yu, Guodong Yang, Liping Zhang, Minyi Guo · 2025

While intelligent applications (e.g., recommendation systems) prefer different CPU-GPU ratios, GPU pooling technique that decouples the GPU and CPU resources yields substantial flexibility when serving diverse applications. With such architecture, DNN-based recommendation services often offload the compute-intensive neural network layers to the remote GPU pool for high resource utilization. However, such a paradigm results in the long end-to-end latency due to two causes: 1) the intermediate data is copied for multiple times during the entire process in current GPU pooling practices, incurring heavy overheads; 2) the content transferred to the GPU pool involves multiple small tensors, suffering from poor bandwidth efficiency. To solve these problems, we design Zero, a runtime system that incorporates a zero-copy transmission mechanism as well as a dynamic tensor merging policy. The zero-copy transmission mechanism unifies memory management across the inference framework and the RPC framework, accompanied by an elaborated serialization protocol to fully eliminate redundant data copying. Meanwhile, the tensor merging policy deliberately organizes small tensors into larger data blocks, so as to transfer them with higher efficiency. Experimental results show that, compared with prior work, Zero reduces the latency of typical recommendation models by up to 15.1% (10.1% on average).

Read the paper · More papers on PaperTik