Enhancing RAG Pipeline Performance with Translation-Based Embedding Strategies for Non-English Documents

Can İşcan, Muhammet Furkan Özara, Akhan Akbulut · 2024

Despite the advances in multilingual embedding models, they often underperform compared to embedding doc-uments in English, which limits the effectiveness of Retrieval-Augmented Generation (RAG) systems in non-English question-answering contexts. This disparity poses a significant barrier for non-English speakers to access high-quality AI-driven re-trieval systems. To address this challenge, this paper proposes a novel approach that leverages English embeddings enhanced by translation. Our approach involves translating texts into English and then embedding them while preserving the original information as metadata. By utilizing robust English and multi-lingual embedding models, we achieve superior results in RAG systems. We present a comprehensive approach for constructing a question-answering system using GPT-40 technology, which leverages translation models to enhance English embeddings. We conducted actual studies to compare the suggested methodology with conventional methods using various embedding models, employing the RAGAS framework. Initial findings indicate that our method greatly improves the retrieval effectiveness of RAG systems in non-English environments, attaining higher context precision and recall metrics. This adaptable solution incorporates English embeddings into multilingual applications, providing a versatile and effective method to improve non-English scenarios by using the capabilities of translation models.

Read the paper · More papers on PaperTik