Performance Efficient Layer-aware DNN Inference Task Scheduling in GPU Cluster

Hongmin Geng, Deze Zeng, Yuepeng Li · GLOBECOM 2022 - 2022 IEEE Global Communications Conference · 2022

GPU has been widely applied to accelerate the DNN based applications. However, single GPU is overwhelmed by the increasing computation requirement of large-scale DNN inference task. Although GPUs cluster alleviates the pressure of massive inference task, it still traps into low efficiency owing to the limited computing power and network bandwidth. Besides, the exclusiveness of GPU device may cause the straggler problem and hence long inference time. Reinforcement Learning (RL) has been widely used in such task scheduling problems. But the delayed reward during the training may slow the convergence speed or even result in non-convergence. To this end, in this paper, we design an improved reinforcement learning based algorithm, called DRM-DQL, to achieve a layer-aware DNN inference task scheduling. We first analyze and model the layer-wise inference task scheduling problem by deep Q-learning. Then, a delayed reward matching strategy is proposed for matching the global reward value to the immediate reward value, which help the algorithm to get the right experience in DNN layer scheduling. The experiment results demonstrate that our algorithm performs better than both heuristic algorithm and the vanilla DQL algorithm, and show the robustness in various network bandwidths, computing power, and DNN model structures.

Read the paper · More papers on PaperTik