Evaluating RAG Pipeline in Multimodal LLM-based Question Answering Systems
Madhuri Barochiya, Pratishtha Makhijani, Hetul Niteshbhai Patel, Parth Goel, Bankim Patel · 2024
Recent improvements have increased the adoption of Multimodal Large Language Models (MLLMs) over traditional uni-modal systems since they can manage and combine information from several types of data. However, current multimodal systems have certain limitations, particularly when it comes to retrieving relevant information for domain-specific queries. This study investigates how Retrieval-Augmented Generation (RAG) approaches can help multimodal LLMs provide more contextually accurate answers from external data sources for Question and Answering (Q&A) systems. The research compares two state-of-the-art models—Google's Gemini-1.0-Pro and OpenAI's GPT-4o-mini for a Q&A system that utilizes Multimodal RAG architecture. Additionally, the study incorporates several embedding models, including Text-embedding-ada-002-v2 and embedding-001, to further improve retrieval performance. In the absence of any standardized criterion, the study used a custom-made multimodal dataset and evaluated the system using six metrics defined by human analysis. The results revealed that using RAG with multimodal LLMs greatly improves performance in Q&A tasks, with GPT-4o-mini slightly outperforming Gemini-1.0-Pro by 5%. These findings promote advancements for research in the field of Multimodal-RAG systems and highlight their potential for much better information retrieval in specialized areas.