Fine-Grained Feature Extraction from Indoor Data to Enhance Visual Question Answering

Rajat Subraya Gaonkar, V A Pruthvi, L Prem Kumar, Rohan Madan Ghodake, M J Raghavendra, Beata Krupa · 2023

Current Visual Question Answering (VQA) models have limited contribution in the area of indoor images. In this paper, the proposed VQA model extracts image and question features from pre-trained VGG16 or ResNetl52v2 and Glove models respectively. Then these features are fed to the novel approach proposed in this paper where Stacked-Attention Network (SAN) which includes customized attention layers with Self-Focus Network (SFN) is used to extract fine-grained features of both the image and question. With these features, the answer prediction model generates an answer. This model is trained and tested on a subset of VQA v2 dataset. It gives a comparable performance when compared with other existing models.

Read the paper · More papers on PaperTik