Dynamic Model Deployment, Batch Scheduling, and Resource Allocation in MLLM-Enabled Edge–Cloud Networks: A Multiagent Two-Timescale DRL Approach
Hualong Huang, Yongkang Du, Wenhan Zhan, Hancong Duan, Kai Peng, Yamin Cheng, Yalan Ye, Zitian Zhao · IEEE Internet of Things Journal · 2025
The deployment of multimodal large language models (MLLMs) on resource-constrained mobile devices poses significant challenges due to their high computational demands. This paper introduces a novel two-timescale optimization framework for efficient MLLM inference in Edge-Cloud networks, addressing the problem of multi-timescale resource management by jointly optimizing slow-timescale MLLMs deployment decisions and fast-timescale batch scheduling, GPU resource allocation, and bandwidth allocation under dynamic network conditions and spatiotemporal request heterogeneity. Our key innovation is a hierarchical twin delayed deep deterministic policy gradient (HALTD3) algorithm that integrates attention mechanisms and long short-term memory networks to optimize slow-timescale MLLMs deployment and fast-timescale resource allocation, minimizing weighted system costs including deployment cost, end-to-end latency, and energy consumption, while meeting stringent quality-of-service requirements. Extensive experiments demonstrate that the HALTD3 algorithm substantially outperforms baseline methods in reducing system costs across diverse MLLM workloads and dynamic network scenarios, validating its effectiveness for practical edge-cloud collaborative inference.