Aggregated Co-attention based Visual Question Answering
Aakansha Mishra, Ashish Prabhu Anand, Prithwijit Guha · 2023
Recent developments in the field of Visual Question Answering (VQA) have witnessed promising improvements in performance through contributions in attention based networks. Most such approaches have focused on unidirectional attention that leverage over attention from textual domain (question) on visual space. This work proposes a multistage co-attention framework. Here, co-attention framework performs both image and text attention. The co-attention mechanism is repeated in multiple stages. Attention on different stages may capture some significant and distinct features for learning better contextual information. Thus, aggregation of attention is performed to preserve the information from different stages. The proposed architecture with multiple stage network could suffer from vanishing or exploding gradients. To prevent this, loss at each of the different stages is computed. Extensive experiments and analysis are performed for validating the effects of aggregated attention and stage-wise loss.