A Bandit-Based Approach to Scheduling LLM Requests in Cloud-Edge Systems
Yandi Li, Jianxiong Guo · 2025
The increasing demand for large language models (LLMs) has placed significant strain on cloud-edge infrastructures, where distributed edges provide low-latency processing but limited resources, and centralized clouds offer scalability at the cost of higher latency. LLM requests, with highly variable processing times, present a unique challenge for scheduling, as their execution times cannot be reliably predicted in advance. In this paper, we propose a novel scheduling framework that models processing time estimation as a contextual combinatorial bandit problem, using a contextual neural upper confidence bound method. The framework dynamically prioritizes LLM requests based on estimated processing times, offloading longer tasks to the cloud and processing shorter tasks at the edge. To further enhance real-world scheduling efficiency and system responsiveness, we incorporate an asynchronous feedback mechanism for model retraining, ensuring that critical scheduling decisions are never delayed by background training updates. Theoretical guarantees of sublinear regret and preliminary experimental results validate the effectiveness of our approach.