A Deep Learning-Based Bengali Visual Question Answering System Using Contrastive Loss
Md. Tareq Zaman, Md. Yousuf Zaman, Faisal Muhammad Shah, Eusha Ahmed Mahi · 2024
Visual Question Answering (VQA) is an interdis-ciplinary research area that uses image recognition, natural language processing, and cognitive understanding to interpret visual data and answer appropriate queries. This study aims to advance Bengali visual question answering (VQA) by proposing a contrastive loss approach to improve bengali VQA models contextual understanding between images and questions. We introduce the new VQA Bengali 1.0 dataset containing 1,864 social media images with 3,728 corresponding yes/no questions, which are splitted into train/test/val sets. Our approach involves extracting image features using ResNet50 and textual features with BangIa BERT variants. These features are fused through an attention mechanism. We incorporate contrastive loss during training to increase discrimination between positive and negative image-text pairs. Comparing combinations of ResNet50 with Sagor Sarker BangIa BERT and Sahaj BangIa BERT models, experiments demonstrate contrastive loss boosts accuracy over cross-entropy loss. The best model with ResNet50, Sahaj BangIa BERT and contrastive loss achieves 93.21 % accuracy. This work advances Bengali VQA through proposing an effective contrastive loss methodology and releasing a balanced dataset. By enhancing model contextual understanding, it has significant potential for further research innovations.