RealArch: A Real-Time Scheduler for Mapping Multi-Tenant DNNs on Multi-Core Accelerators

Xuhang Wang, Zhuoran Song, Xiaoyao Liang · 2023

Nowadays, the significance of multi-tenant deep neural networks (DNNs) has grown exponentially, particularly for cloud providers who execute multiple DNN models on one server to fulfill the users’ requirements while reducing the computational overhead. To satisfy the heavy computation requirement of multi-tenant DNNs, a feasible approach is to establish a multi-core accelerator housing multiple sub-accelerators. Although many researchers have achieved a certain success by designing either offline schedulers for heterogeneous accelerators or real-time schedulers for homogeneous accelerators, they fail to schedule multi-tenant DNNs to both homogeneous and heterogeneous multi-core accelerators in real time given the large search space and restricted overhead constraint.In this paper, we propose RealArch, a novel real-time scheduler that efficiently schedules multi-tenant DNNs to both homogeneous and heterogeneous multi-core accelerators in real-time. The key idea of RealArch is to quickly find the minimal latency of mapping multi-tenant DNNs to sub-accelerators, considering the occupation of DRAM, sub-accelerators, and buffers. To support the key idea, we first establish lightweight estimation models for multiple sub-accelerators to evaluate the Data Movement (DM) and Execution (EX) time when mapping a layer to them. Then, we design a real-time scheduling algorithm to compute and select the mapping solution with minimal latency. Finally, we build a low-cost hardware scheduler to perform the estimation models and the real-time scheduling algorithm. Extensive experiment results verify that RealArch can exceed the baseline Round Robin scheduling algorithm and two state-of-the-art schedulers AI-MT and MAGMA with acceptable hardware overhead.

Read the paper · More papers on PaperTik