The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Matías Martínez · 2025
Open-source large language models (LLMs) enable developers to create AI-based solutions while maintaining control over aspects such as privacy and compliance, thereby providing governance and ownership of the model deployment process. To utilize these LLMs, inference engines are needed. These engines load the model's weights onto available resources, such as GPUs, and process queries to generate responses. The speed of inference, or performance, of the LLM is critical for real-time applications. In this paper, we analyze the performance, particularly throughput (tokens generated per unit of time), of 20 LLMs using two inference libraries: vLLM and HuggingFace's pipelines. We investigate how various hyper-parameters, which developers must configure, influence inference performance. Our results reveal the importance of hyperparameter optimization to achieve maximum performance. We demonstrate that applying hyperparameter optimization during GPU upgrades or downgrades can enhance throughput. For instance, in Hugging-Face pipelines, this approach yields average improvements of 9.16% and 13.7%, respectively.