JoAt: to dynamically aggregate visual queries in transformer for visual question answering

Mingben Wang, Juan Yang, Lixia Xue, Ronggui Wang · 2024

Attention mechanisms have shown impressive abilities in solving downstream multi-modal tasks. However, there exists a natural semantic gap between vision and language modalities that hinders conventional attention-based models in achieving effective cross-modal semantic alignment. In this paper, we present JoAt, a Joint Attention net, through which we investigate how to utilize the visual background information more directly in a query-adaptive manner to enrich querying semantics for each visual token, and how to more fully bridge the semantic gap to achieve cross-modal alignment between visual-grid and textual features. Specifically, our JoAt utilizes each query’s neighboring pixels, aggregates the visual query tokens from different receptive fields, and allows the model to dynamically select the most relevant neighboring tokens for each query, then obtains representations that are more semantically matched with the textual features to realize better interaction between visual and linguistic modalities. The experimental results show that our JoAt net can fully utilize different semantic-level signals from visual features at different receptive fields and effectively narrow the natural semantic differences between visual and language modalities. Our JoAt achieved an accuracy of 72.15% and 98.90% on the VQAv2.0 test-std and CLEVR benchmarks, respectively.

Read the paper · More papers on PaperTik