Vision Question Answering System Based on Roberta and Vit Model
Haiyang Tang · 2022 International Conference on Image Processing, Computer Vision and Machine Learning (ICICML) · 2022
With the advent of the information age, VQA has become a key research direction cross the field of CV and NLP. Based on a large number of real scene images on the KAGGLE platform, this paper applies the Transformers model to the CV field, and combines the Transformers based NLP algorithm to establish a VQA system. The experimental results show that our model can give accurate answers in a simple and orderly scenario, and there is a certain deviation between the generated results and the real answers in a chaotic scenario, which verifies the effectiveness of our model. Wup similarity is selected as a metric in this paper. The results show that the RoBERTa-ViT model has the most outstanding performance in VQA task, with the metric value of 0.351. Surprisingly, the BERT-SwinT model performs relatively poorly. which may be because the SwinT model is more complex than the VIT model and does not show its own advantages in small data set task.