Integrating Dynamic Routing with Reinforcement Learning and Multimodal Techniques for Visual Question Answering

Xinyang Zhao, Zongwen Bai, Meili Zhou, Xincheng Ren, Yuqing Wang, Linchun Wang · 2024

The critical task of Visual Question Answering (VQA) is how the model can effectively process and understand the different levels of visual information in an image and closely associate the visual information with the keywords in the text. This technique is the key to enhancing VQA. Most of the successful attention mechanisms are mainly built on top of shallow models, while the deep Transformer architecture to achieve effective fusion of global and local dependencies is still an open problem, and optimizing its Transformer architecture for VQA systems is crucial. Therefore, textual research focuses on how to process textual information while deeply parsing image visual elements. This study proposes the IRM model (Integrating Dynamic Routing with Reinforcement Learning and Multimodal Techniques) by improving the traditional Transformer architecture. Moreover, the GSA module (Gates Self-Attention) filters essential information and enhances the model’s ability to perceive location information. In addition, the MIL module (Multimodal et al.) is proposed to enhance the deeper interaction between text and image information. To confirm the effectiveness of these improvements, tests were conducted on the VQA-v2 dataset, and the experimental results show that this method outperforms some existing state-of-the-art techniques in terms of accuracy.

Read the paper · More papers on PaperTik