Similarity Thresholds in Retrieval-Augmented Generation

Irina Radeva, Ivan P. Popchev, Miroslava Dimitrova · 2024

This paper evaluates the performance of open-source Large Language Models (LLMs) Mistral: 7b, Llama2:7b, and Orca2:7b within the context of Retrieval-Augmented Generation (RAG). The purpose of the study is to identify the similarity score thresholds that yield the best performance across Natural Language Processing metrics. Despite the advancements in LLMs, their ability to access up-to-date, domain-specific information remains limited. RAG addresses this by integrating external database querying to produce contextually comprehensive and precise responses. This study investigates the similarity score thresholds for the LLMs using a custom-developed dataset from the domain of agriculture. Tests were conducted using the PaSSER web-based application, which was enhanced to support detailed configuration of retrieval processes. The PaSSER App is a complementary project to the Smart Crop Production Data Exchange (SCPDx) platform. The results showed that Mistral performs best at lower similarity threshold (0.55), Orca achieves the highest performance at a moderate threshold (0.65), and Llama at a lower threshold (0.55), but has strong performance at higher thresholds as well (0.70 and 0.85). The study discusses several factors that could cause the differences in the similarity score thresholds of the assessed LLMs. Future development will focus on leveraging other pre-trained open source LLMs (over 40b), exploring approaches to fine-tuning, and further integration on PaSSER App with the existing infrastructure of the SCPDx platform.

Read the paper · More papers on PaperTik