LLMConf: Knowledge-Enhanced Configuration Optimization for Large Language Model Inference

Jingkai He, Pengfei Chen, Yilun Wang, Haiyu Huang, Chuanfu Zhang, Haojia Huang, Danwen Chen · 2025

As large language models (LLMs) are widely applied across various domains, improving the quality of LLM inference services is essential. In this paper, we find that optimizing configuration parameters of LLM inference engines can significantly improve LLM inference performance in terms of latency and throughput. Therefore, we propose LLMConf, an automated performance tuning system that optimizes multiple LLM inference performance metrics by searching for the optimal configuration parameters of the LLM inference engine. We first introduce a knowledge-enhanced approach to identify the set of configuration parameters (LLMConfigs) that most significantly impact LLM performance from the adjustable parameters provided by the LLM inference engine. We then perform automated data collection to build functional relationships between LLMConfigs and each performance metric. Additionally, LLMConf employs a multi-objective optimization module to obtain optimal LLMConfigs for simultaneously optimizing multiple performance metrics. The experimental results show that LLMConf significantly outperforms existing methods. Compared to the default configuration parameters of the LLM inference engine, LLMConf achieves an average improvement of 20.1% across 7 key performance metrics. Moreover, experiments demonstrate that LLMConf has strong transferability across diverse datasets, varying concurrency levels and different LLM base models.

Read the paper · More papers on PaperTik