Multimodal Question Answering with DenseNet and BERT for Improved User Interaction

Battula Bhavana, Chennu Chaitanya, Bharathi Mohan G · 2025

Visual Question Answering (VQA) is a challenging AI task that combines computer vision and natural language processing to generate contextual answers from image inputs. It requires a model that is capable of comprehending both visual and textual information, generating correct reasoning and context-based predictions. The proposed model uses DenseNet to capture deep visual features and BERT to process questions so that image-text correlation can be interpreted more judiciously. With the ability to handle complex images and different linguistic structures, this system is most appropriate for education, healthcare, security, and accessibility applications, in which automatic question answering based on images enhances decision-making and user participation.Existing approaches, such as machine learning-based question generation, attention model, and Faster R-CNN-based methods, have drawbacks like inconsistent responses, high computational overhead, and scalability. The introduced model addresses the feature extraction process, increases contextual alignment, and maximizes training efficiency with learning rate scheduling and mixed-precision training, and it is a more efficient and scalable method than the current techniques.

Read the paper · More papers on PaperTik