Transformer-Based Approaches for Multilingual Visual Question Answering
Linh Xuan Truong, Vinh Quang Pham, Kiet Van Nguyen · International Journal of Asian Language Processing · 2022
Visual Question Answering (VQA) is a difficult task that has steadily gained popularity and made significant progress. VQA is also one of the potential tasks with a combination of Computer Vision and Computational Linguistics. Through the prism of the English language, VQA has mostly been researched. But using the same strategy to VQA in other languages would demand a significant investment in resources. Therefore, finding solutions to the multilingual VQA (mVQA) task is crucial. This topic is very challenging because it requires the model not only to comprehend the visual context of the image but also to understand the language and provide the correct answer to the input question in that language. Utilizing the outstanding development and remarkable performance of the Vision-Language Pre-training models, this paper presents our approach to the Multilingual Visual Question Answering shared tasks (EVJVQA) ( https://vlsp.org.vn/vlsp2022/eval/evjvqa ) of the Vietnamese Language and Speech Processing (VLSP) Challenge 2022. We propose a transformer-based framework with a specialized custom encoder consisting of two modules for question–answer embedding and visual embedding utilizing the Vision-Language Pre-training models. Without using any data augmentation techniques, our strategy performed better than the organizer’s baseline model and produced results that were competitive with those of other approaches on the public and private tests of the UIT-EVJVQA benchmark. Our highest result on the private test set is 34.67% and 41.15% in the two metrics BLEU and F1-score, respectively.