SPLIT: QoS-Aware DNN Inference on Shared GPU via Evenly-Sized Model Splitting
Diaohan Luo, Yu Tian, Yuewen Wu, Heng Wu, Tao Wang, Wenbo Zhang · 2023
Improving QoS by simultaneously reducing the latency violation rate and jitter in the presence of multiple deep learning inference (DLI) tasks sharing a single edge computing processor remains a challenge. However, existing DLI systems at the edge, designed to maximize throughput, face performance challenges when confronted with requests with varying QoS.