M3-RAG: Unified Multimodal and Multilingual Retrieval-Augmented Generation

Junqi Xu, Junjie Zhu, Yaohan Ba · 2025

Retrieval-Augmented Generation (RAG) systems grapple with inherent limitations in processing multilingual and multimodal queries, particularly when bridging linguistic diversity with cross-modal semantic alignment. To address this, we present M3-RAG, an integrated framework that synergizes multilingual text retrieval (via BGE-M3), cross-lingual image retrieval (via AltCLIP), and language-contextualized generation within a unified architecture. At its core, the framework introduces three pivotal advancements: First, a zero-shot alignment mechanism employing unsupervised contrastive learning bridges text and image embeddings across 12 languages, achieving 83.5% cross-modal retrieval accuracy a 12.3 % absolute improvement over existing state of the art methods. Second, a dynamic modality router with lightweight MLP architecture dynamically activates text/image retrieval paths based on query semantics, attaining 92.3 % routing accuracy while reducing redundant computations by 28 % through adaptive threshold optimization. Third, we construct the Cross-MMQA benchmark, the first multilingual multimodal QA dataset encompassing$\mathbf{1 5 K}$culturally annotated text-image pairs across 12 languages, rigorously validated through crowdsourced bilingual annotations. Comprehensive evaluations demonstrate M3-RAG's superiority, achieving 72.1 % F1 on mixed-modal queries-surpassing textonly and image-only baselines by 22.4 % and 15.8 % respectively while maintaining real-time responsiveness (587 ms /query) on consumer-grade GPUs through FP16 quantization and query caching. The system's efficacy is further validated in real-world deployments, where it improved diagnostic accuracy by 41 % for Swahili-speaking medical consultations compared to conventional RAG systems.

Read the paper · More papers on PaperTik