RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering

Yang Bai, Christan Grant, Daisy Zhe Wang · 2025

Multi-modal retrieval-augmented Question Answering (MRAQA), integrating text and images, has gained significant attention in information retrieval (IR) and natural language processing (NLP).Traditional ranking methods rely on small encoder-based language models, which are incompatible with modern decoder-based generative large language models (LLMs) that have advanced various NLP tasks.To bridge this gap, we propose RAMQA, a unified framework combining learning-torank methods with generative permutationenhanced ranking techniques.We first train a pointwise multi-modal ranker using LLaVA as the backbone.Then, we apply instruction tuning to train a LLaMA model for re-ranking the top-k documents using an innovative autoregressive multi-task learning approach.Our generative ranking model generates re-ranked document IDs and specific answers from document candidates in various permutations.Experiments on two MRAQA benchmarks, We-bQA and MultiModalQA, show significant improvements over strong baselines, highlighting the effectiveness of our approach.

Read the paper · More papers on PaperTik