Optimizing Confidence Scoring in RAG-Based LLM Chatbots for Technical Support Services: A Prompt Engineering Approach

W.T. Chan, Kevin Hung, Raymond H. Ho, Gary Man-Tat Man · 2025

The global shortage of skilled personnel in technical support has spurred significant interest in AI-driven solutions, particularly Large Language Model (LLM)-based customer service chatbots. However, a critical challenge in deploying these systems lies in addressing AI hallucination, wherein models generate responses that are plausible yet factually incorrect. This study investigates a prompt engineering approach to enhance confidence estimation and mitigate AI hallucinations in LLM chatbots. Three distinct prompt strategies—Basic, Advanced, and Combo prompts–are systematically evaluated to improve response reliability. Given that LLMs inherently lack the ability to explicitly express uncertainty (e.g., by stating “I don't know”), a structured confidence scoring mechanism is employed to refine accuracy and reduce the Expected Calibration Error (ECE). Experimental results reveal that Basic prompts achieve an accuracy of 69.33% (ECE: 23.33), Advanced prompts improve accuracy to 75.33% (ECE: 14.87), and Combo prompts further enhance accuracy to 81.33% while reducing ECE to 8.4. These findings underscore the efficacy of prompt engineering in mitigating AI hallucinations and advancing the performance of LLM chatbots in real-world customer support applications.

Read the paper · More papers on PaperTik