Performance Evaluation for Cost-Effective Retrieval Process for Multi-Document Retrieval-Augmented Generation on a Domain-Specific Dataset
Maggiore Alica Kartiyanta, Eugenia Ancilla, Kenny Jingga · 2025
Retrieval-Augmented Generation (RAG) has emerged as a promising approach to enhance Large Language Models (LLMs) in domain-specific tasks without requiring expensive fine-tuning. While paid dense embedding models, such as OpenAI’s text-embedding-3, are widely used in retrieval systems, their high cost raises concerns about whether free or traditional retrieval methods can offer comparable performance. This study evaluates the effectiveness of costeffective sparse and dense retrieval approaches within a RAG framework using the RAGEval DragonBall dataset, focusing on multi-document retrieval in the medical domain. Final retrieval evaluation using GPT-40-mini as the generator, with RAGAs and BERTScore as evaluation metrics, demonstrates that the traditional keyword-based retrieval method, BM25 outperforms others in several aspects: context recall (1.0), answer relevance (0.9087), factual correctness (0.7122), context entity recall (0.7741), semantic similarity (0.9478), and BERTScore (0.9277). Additionally, a hybrid model combining BM25 with dense embedding model and chunking yields competitive results, achieving the highest factual correctness (0.7789), semantic similarity (0.9515), and answer relevance (0.9094), while also maintaining perfect context recall (1.0) and the lowest noise sensitivity (0.1111), showing the possibility of combining both retrieval methods for improved performance. These results challenge the assumption that paid embeddings are always the better option and highlight the potential of costeffective retrieval strategies for domain-specific applications.