Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models
Mu Yuan, Lan Zhang, Di Duan, Liekang Zeng, Miao-Hui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li · IEEE Transactions on Mobile Computing · 2025
Serving machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches.