Optimizing Retrieval-Augmented Generation Chatbot with Hyperparameter Tuning
Novem Ardan Rohmadin, Ridi Ferdiana, Indriana Hidayah · 2025
This study presents the design, implementation, and evaluation of a Retrieval-Augmented Generation (RAG)-based question-answering (QA) chatbot enhanced with Hyperparameter Optimization (HPO). As QA systems grow in importance for delivering timely and accurate information, optimizing performance becomes essential. The proposed system tunes key RAG parameters, such as embedding model, chunk size, chunk overlap, top-k, temperature, and reranker threshold. This experiment involved testing the following methods: Grid Search, Random Search, Upper Confidence Bound (UCB) and Bayesian Optimization. Five core evaluation metrics are used: faithfulness, answer relevancy, context recall, latency, and token usage. These are scalarized into a single objective function to balance response quality and efficiency. Bayesian Optimization achieved the best scalarized score (0.314), with UCB performing comparably. Further analysis reveals a tradeoff between latency and quality. Low-latency configurations (under 1.7 seconds) often result in reduced faithfulness and context precision, while moderate-latency setups (1.6–2.2 seconds), especially from UCB and Bayesian methods, achieve higher composite Quality Scores. Although increased top-k improves recall, it also adds latency. Answer relevancy remains consistently high across methods. Overall, integrating HPO into RAG pipelines enhances QA performance and adaptability, offering a reproducible approach to building efficient, latency-aware systems suitable for real-world deployment. This approach is especially valuable for chatbot applications where both response accuracy and timeliness are critical.