Evaluating Accuracy in Large Language Models: Benchmarking Corrective Rag Vs. Naive Retrieval Augmented Generation Approach

Rajendra Gangavarapu, Aswath Ram Adayapalam Srinivasan, Venkata Moparthi · 2025

Retrieval-Augmented Generation (RAG) has emerged as a promising approach to mitigate the limitations of Large Language Models (LLMs) in generating factually accurate and consistent text. The main focus of this technical survey is on correct answers as the key performance indicator (KPI) for comparing two well-known RAG methods: Naive RAG and Corrective retrieval augmented Generation (CRAG). Naive RAG exhibits a strong dependence on the relevance of retrieved documents, resulting in suboptimal performance when retrieval quality is compromised. CRAG, on the other hand, adds new features to improve robustness and adaptability, such as a retrieval evaluator, large-scale web searches, and a decompose-then-recompose algorithm. We introduce the Comprehensive RAG Benchmark (CRAG), which encompasses a diverse set of question-answer pairs spanning multiple domains, categories, entity popularities, and temporal dynamics, to facilitate a comprehensive evaluation of RAG models' performance in generating correct answers. Experiments show that adding RAG to LLMs makes a big difference in how many correct answers you get, with CRAG consistently beating Naive RAG. Nevertheless, RAG models, including CRAG, demonstrate lower answer correctness when confronted with questions pertaining to highly dynamic, less popular, or more complex facts. These results make it clear that more research and development is needed to make RAG models more reliable and able to give correct answers in a wider range of situations.

Read the paper · More papers on PaperTik