Text-based Visual Question Answering Based on Text-Aware Pre-Training
Yupeng Zhang, Jianhua Wu, Zhengkui Chen, Hai Huang, Guide Zhu · 2024
Text-based Visual Question Answering (TextVQA) is a subfield of Visual Question Answering (VQA) that is able to read the text in a given image. Existing work on TextVQA usually improves model performance by designing more powerful model architectures. However, the performance improvements of these custom methods are usually not significant. The text-aware pre-training method can enable the model to better understand the relationship between questions, scene text and visual objects in the image, thereby providing a new way to improve the performance of existing custom text-based visual question answering methods. In this paper, we propose a TextVQA method called TAP+that improves the state-of-the-art text-aware pre-training (TAP) method in two aspects. First, a vision-language pretraining model VisualBERT is adopted as the base backbone network for multimodal fusion, significantly improving the accuracy of the model. Second, an attention module is introduced to filter out redundant or irrelevant features output from the fusion module, thereby further improving the accuracy of the model. Experimental results show that our model outperforms TAP model and other non-pretraining baselines, achieving an accuracy of 49.81% on the TextVQA dataset test set and an ANLS score of 0.569 on the STVQA dataset test set.