RFPG: Question-Answering from Low-Resource Language (Arabic) Texts using Factually Aware RAG

Mitha Alshammary, Md Nahiyan Uddin, Latifur Khan · 2024

Since Large Language Models (LLMs) face several challenges, including hallucinations, the Retrieval-Augmented Generation (RAG) model has been proposed as a solution. While RAG has been widely used for English, it remains underexplored in low-resource languages like Arabic. This study enhances the RAG model by focusing on a specific domain within Arabic, a low-resource language. Our proposed framework, RFPG, addresses question-answering by integrating (a) fact-checking into the retrieval process and (b) customized and innovative prompts. Our model was tested on 123 questions and it was able to answer with an accuracy of 100% and reference the sources with a precision of 98%, outperforming both RAG and standard LLMs, including the latest models like GPT-4o, GPT-4o mini, and GPT-4, in Arabic text question-answering.

Read the paper · More papers on PaperTik