CMACC: Cross-Modal Adversarial Contrastive Learning in Visual Question Answering Based on Co-Attention Network
Zepu Yi, Songfeng Lu, Xueming Tang, Junjun Wu, Zhu Jian-xin · 2024
Visual Question Answering (VQA) as a cutting-edge domain blending computer vision and natural language processing has garnered significant research momentum. Nev-ertheless, latest VQA frameworks which leverage CNN and RNN technologies to extract features and delve into the multimodal interplay between text and images, encounter difficulties in integrating local features with global dependencies. This integration challenge hinders the models' ability to precisely grasp the pivotal aspects of the answer. Furthermore, the dearth of training data poses a significant impediment to models' effective learning of multimodal information exchange and problem semantics. To address these issues, our model introduces a new method called CMACC (Cross-Modal Adversarial Contrastive Learning Based on Co-Attention). Through cross-modal adversarial contrastive learning, combined with common attention, image and text information are integrated. Adversar-ial learning aligns the latent feature distribution between text and images, and contrastive learning aligns multimodal sample features in the same context to enhance the adaptability of the model. In addition, advanced data augmentation techniques are integrated to further enhance the model's adaptability to different scenarios and problem types. We conducted experimental evaluations on three widely used VQA datasets (VQA v1.0, VQA v2.0, and COCO-QA) and the results showed that CMACC has significant improvements in accuracy and generalization performance compared to traditional methods.