Two-Sided Matching for Batch-Aware LLM Request Scheduling in Edge Networks

Tan Li, Yuhang Gong · 2025

Edge Large Language Model (LLM) inference offers reduced latency and enhanced privacy compared to cloud-based approaches. Current inference engines utilize batching mechanisms for computational efficiency. However, static batching creates prompt interdependencies, leading to inefficient memory usage and prolonged processing times due to batch heterogeneity. Existing request scheduling solutions optimize single-server batching but cannot coordinate across multiple edge servers or handle resource constraints at edge. In this paper, we propose a batch-aware request scheduling framework formulated as a two-sided matching game, where batch composition affects individual inference performance through peer effects. We design novel utility functions for users and edge servers based on inference quality, memory and computing capability, construct preference lists, and establish stable assignments via Gale-Shapley algorithm. Based on this, beneficial swaps further optimize batch homogeneity while preserving service quality. Extensive simulations demonstrate substantial improvements in service quality, consistently outperforming baselines across varying request volumes and resource constraints.

Read the paper · More papers on PaperTik