Energy Efficient GPU Frequency Scaling Policy for Inference Serving using Queue Model
Yuhao Zhang, Wenqian Zhang, Xiaowen Huang, Guanglin Zhang · 2025
The rapid advancement of intelligent applications driven by large language models (LLMs) has significantly increased computational demands on cloud/edge computing infrastructures. These centers are commonly equipped with Graphics Processing Unit (GPU) to meet the high computational demands of LLMs. However, the substantial power consumption of GPUs poses a critical challenge in optimizing data center energy consumption without compromising task latency. In this paper, we propose QEEFS, a Queue-theory-based Energy-Efficient Frequency Scaling policy to address the trade-off between latency and energy consumption in LLM inference processes. Specifically, we model the LLM inference server as a M/G(r, f)/1/N queue, capturing the stochastic nature of inference requests and predicting average service time under various server configurations. QEEFS optimizes GPU operating frequencies through offline profiling and online probing phases, making use of DVFS technology to dynamically control the energy consumption of individual GPUs. Evaluation on an NVIDIA RTX V100 Tesla GPU demonstrates the efficacy of our approach in reducing energy consumption while maintaining acceptable latency.