Cloud-Edge System for Scheduling Unpredictable LLM Requests With Combinatorial Bandit
Yandi Li, Jianxiong Guo, Zhiqing Tang, Xingjian Ding, Juncheng Wang, Tian Wang, Weijia Jia · IEEE Transactions on Services Computing · 2025
The rapid growth in demand for large language models (LLMs) has strained cloud-edge infrastructure. While edges offer low latency and clouds provide vast resources, scheduling LLM requests efficiently remains a major challenge due to their unpredictable processing times, which leads to Headof-Line (HOL) blocking that degrades system throughput and responsiveness. To address this, we introduce the Online CloudEdge Collaborative Request Scheduling (OCE-CRS) framework. OCE-CRS models the proactive scheduling of LLM requests as a contextual combinatorial bandit problem. At its core is our novel Combinatorial Neural Delayed Upper Confidence Bound (CN DUCB) algorithm, which learns to predict request processing times from the semantic content of the request prompt alone. This enables an inspired policy based on Shortest Job First (SJF) that prioritizes shorter jobs for edge execution, simultaneously maximizing throughput and mitigating HOL blocking. To prevent time-consuming neural network training from blocking scheduling decisions, we employ an asynchronous mechanism. This decouples model updates from the real-time scheduling loop, effectively handling the resultant delayed feedback where observations from past rounds are used in later training steps. We provide a theoretical sublinear regret bound for our algorithm. Extensive experiments validate that OCE-CRS significantly improves throughput, Job Completion Time (JCT), and queueing delay, demonstrating superior performance and robustness in both static and continuous batching environments.