Jily: Cost-Aware AutoScaling of Heterogeneous GPU for DNN Inference in Public Cloud

Wang Zhaoxing, Xuehai Tang, Qiuyang Liu, Jizhong Han · 2019

Recently, a large number of DNN inference services have emerged in public clouds, making the low-cost deployment of DNN inference services a hot research topic. Previous studies have failed to take into account GPU heterogeneity and batch processing, both of which will seriously affect the financial cost as well as the latency. In this paper, we study the problem of DNN inference service deployment in public cloud, considering both GPU heterogeneity and batch processing. The goal is to minimize the financial costs under the constraint of latency. We propose Jily, an autoscaling scheduler for DNN inference services to minimize the cost while satisfying the given latency SLO. Jily finds the optimal heterogeneous GPU instance provisioning through a DNN inference model profiler, a latency estimator, a workload predictor and a cost-aware scaler. Simulation results demonstrate that Jily can reduce average cost by up to 28% compared to a state-of-the-art autoscaling approach. Further, Jily has been proved to have good versatility and robustness under different batching mechanisms and latency SLO constrains.

Read the paper · More papers on PaperTik