Integrating External Knowledge with LLMs: A Systematic Review of RAG Approaches

Ivan Mikulić, Marin Vlaić, Goran Delač, Marin Šilić, Klemo Vladimir · 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation, exhibiting exceptional in-context learning abilities without requiring extensive fine-tuning. However, these models often suffer from significant limitations, including hallucinations where fabricated or incorrect information is presented as fact, and the need for constant retraining in order to effectively integrate new knowledge. Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm to address these challenges by combining LLMs with information retrieval (IR) techniques. By leveraging external knowledge bases, RAG pipelines retrieve and incorporate relevant information dynamically, enhancing the factual accuracy and adaptability of LLMs. This paper provides an overview of the state-of-the-art methods for implementing RAG in LLMs, examining key components of the RAG pipeline, including data preparation, retrieval, reranking, and post-retrieval techniques. We aim to highlight how these components collectively enable LLMs to achieve greater reliability and flexibility, paving the way for more robust AI applications.

Read the paper · More papers on PaperTik