Embedding Attention Blocks For Answer Grounding
Seyedalireza Khoshsirat, Chandra Kambhamettu · 2024
Despite the introduction of various attention methods for the answer grounding task, they often encounter three common challenges. Firstly, some designs lack the capability to utilize pre-trained networks and fail to benefit from extensive data pre-training. Secondly, certain custom designs are not based on well-established previous models, limiting the network’s learning potential. Lastly, complex designs that impede reimplementation or enhancement. To address these issues, this paper presents a novel architectural block, termed the Embedding Attention Block (EAB). This block re-calibrates channel-wise image feature-maps by explicitly modeling inter-dependencies between the image feature-maps and the image-question-answer embeddings. The visual demonstration showcases how this block filters out irrelevant feature-map channels based on embeddings. Our approach builds upon three key concepts: Semantic Region Proposal, Input Embedding, and Dynamic Region Fusion. We validate our method’s efficacy using the TextVQA-X, VQS, VQA-X, and VizWiz-VQA-Grounding datasets, and carry out several ablation studies to show the effectiveness of our design choices. Notably, our novel network ranked first place on the 2023 VizWiz-VQA-Grounding challenge leaderboard and won the challenge.