RAG Certainty: Quantifying the Certainty of Context-Based Responses by LLMs

Kento Hasegawa, Seira Hidano, Kazuhide Fukushima · 2024

Large language models (LLMs) have recently been employed for a wide variety of purposes. Retrieval-augmented generation (RAG), in which an LLM generates a response based on context relevant to the prompt, is often used to enable the LLM to adapt to specialized domains. However, sentences generated by a generative LLM may contain incorrect information, known as “hallucinations.” The challenge in identifying hallucinations within the RAG framework involves evaluating the certainty of both context retrieval and LLM outputs. In this paper, we propose a metric called RAG certainty to quantify the certainty of LLM outputs within a RAG framework. The proposed metric is calculated based on certainty scores from both information retrieval and response generation. Experimental results demonstrate that the proposed metric effectively reflects the certainty of information retrieval in a RAG framework. We further validated the proposed metric through a case study that assesses the predicted Common Vulnerability Scoring Sys-tem (CVSS) scores for cybersecurity vulnerabilities and found that errors are mitigated according to the proposed metric.

Read the paper · More papers on PaperTik