MRAN-VQA: Multimodal Recursive Attention Network for Visual Question Answering
Mohammad Shariful Islam, Mohammad Abu Tareq Rony, Md Murad Hossain Sarker, Md. Khairul Bashar Bhuiyan, Md Saib, Md. Aktarujjaman, Md Shahab Uddin, Abeer D. Algarni, Ahmad Taher Azar, Walid El‐Shafai · Engineering Science and Technology an International Journal · 2025
Visual Question Answering (VQA) is a fundamental challenge in multimodal AI, requiring models to integrate and reason over both visual and textual information. Despite advancements in deep learning, existing VQA models struggle with multi-step reasoning, hierarchical feature fusion, and multilingual generalization, limiting their effectiveness in real-world applications. This paper introduces MRAN-VQA, a Multimodal Recursive Attention Network for VQA, designed to address these limitations through a three-stage reasoning pipeline. The proposed approach first employs Recursive Attention Encoding, where a Vision Transformer (ViT) extracts visual features, and BERT-based embeddings encode textual information. A recursive self-attention mechanism iteratively refines these representations, improving contextual alignment. Hierarchical Feature Fusion integrates multi-level visual–text interactions through bilinear attention pooling and gated cross-modal operations. Finally, Answer Prediction with Attention Grounding applies a self-attentive reasoning module to responses while optimizing an Attention Grounding Score (AGS) for improved interpretability. Experiments on VQA v2.0, CLEVR, and our custom BanglaVQA datasets demonstrate that MRAN-VQA outperforms state-of-the-art models, achieving 75.6% accuracy on VQA v2.0, 96.1% on CLEVR, and 72% on BanglaVQA—notably surpassing transformer-based baselines. The model exhibits superior multi-step reasoning capabilities in compositional queries and significantly enhances performance in low-resource multilingual settings. The model extracts textual features via BERT and visual features via a Vision Transformer (ViT). These are refined through a Multimodal Recursive Attention Network (MRAN) using iterative Text-to-Image and Image-to-Text Attention. Attentive fusion combines the modalities, and a softmax layer produces the final prediction. Recursive attention enables deeper contextual understanding, enhancing multimodal reasoning in complex VQA tasks. • MRAN-VQA introduces recursive attention for enhanced multi-step reasoning in VQA. • Proposes hierarchical fusion of low, mid, and high-level multimodal features. • Defines an Attention Grounding Score (AGS) with a trainable auxiliary loss. • Demonstrates strong multilingual and low-resource performance on BanglaVQA. • Outperforms recent SOTA on VQA v2.0, CLEVR, and BanglaVQA (best trade-off at R = 4).