Advancing Visual Question Answering Through Multimodal Fusion and Deep Learning Techniques

Parth Bhatnagar, Manjit Singh Sodhi · 2025

This paper introduces an advanced Visual Question Answering (VQA) approach that leverages multimodal learning to bridge the gap between image and text understanding, aiming for enhanced accuracy and contextual relevance in response generation. The proposed method combines convolutional neural networks (CNNs)—with ResNet50 selected for its superior performance in visual feature extraction—and state-of-the-art language models for sophisticated question parsing. By fusing these visual and textual features, the model effectively addresses complex questions about visual content, achieving high precision and context-aware responses. Empirical analysis reveals that the ResNet50-based model consistently outperforms alternative architectures, achieving significant improvements in accuracy and robustness across varied datasets. Key performance metrics, including accuracy, F1 score, and response consistency, validate the model's efficacy in multimodal comprehension tasks. The findings underline the importance of multimodal integration for VQA applications, showcasing the ResNet50 architecture as a critical component for achieving best-in-class performance. The paper concludes by discussing potential advancements in VQA, including deeper multimodal fusion strategies and the exploration of transformer-based image-language interactions, to further refine model accuracy and adaptability.

Read the paper · More papers on PaperTik