Performance Analysis of LLM Models with RAG and Fine-Tuning T5 for Chatbot Optimization in Call Centers
Lukman Arif Sanjani, Riyanarto Sarno, Kelly Rossa Sungkono, Agus Tri Haryono, Abdullah Faqih Septiyanto, Dwi Sunaryono · 2025
This research investigates the comparative performance of various Large Language Models (LLMs) to determine the most suitable model for XYZ Company, a call center organization aiming to integrate a chatbot solution. Chatbots significantly improve call center operations by streamlining interactions, reducing the workload on human agents, and enhancing the overall customer experience. Unlike many chatbot datasets that are document-based or contain contextual fields, the dataset of XYZ Company is composed of question-answer pairs. This unique data structure requires a customized approach to model selection. Experiments were conducted on multiple models, with RAG using Llama3 emerging as the top performer. Results indicate a BLEU score of 18%, ROUGE-1 of 46%, ROUGE-2 of 27%, ROUGE-L of 36%, BERTScore Precision of 89%, Recall of 91%, F1 of 90%, and METEOR score of 41%. These metrics underscore RAG with Llama3 to deliver accurate, efficient responses, supporting the goal of XYZ Company to implement an effective, real-world chatbot solution.