MediatorDNN: Contention Mitigation for Co-Located DNN Inference Jobs
Seyed Morteza Nabavinejad, Sherief Reda, Tian Guo · 2024
With the increase in computing power of cutting-edge hardware platforms, it is a common practice to run multiple jobs on a single machine for improved resource utilization and throughput. However, this leads to inevitable resource contention among co-located jobs, impacting their performance. The resource contention can worsen due to fluctuations in resource utilization of jobs caused by variations in their input workload. To tackle the co-location contention for DNN inference jobs, we propose MediatorDNN, which considers contention and resource utilization variation when co-locating DNN inference jobs. It profiles each DNN, monitors microarchitectural metrics such as memory bandwidth and cache access pattern, and high-level resource utilization like CPU utilization. Based on profiling results and leveraging Modern Portfolio Theory (MPT), MediatorDNN decides on the co-location of jobs. Experimental results with various DNNs on two hardware platforms show that MediatorDNN improves throughput by up to 108% (21 % on average) compared to an approach only considering contention and ignoring resource utilization variation.