ECNU_MIV at SemEval-2023 Task 1: CTIM - Contrastive Text-Image Model for Multilingual Visual Word Sense Disambiguation
Zhenghui Li, Qi Zhang, Xueyin Xia, Yinxiang Ye, Qi Zhang, Cong Huang · 2023
This paper describes the performance of MIV team in SemEval-2023 Task-1-Visual Word Sense Disambiguation (Visual-WSD).This task is to give a potentially ambiguous word and some limited textual context and select among a set of candidate images the one which corresponds to the intended meaning of the target word, which plays a critical role in human language understanding.Our team focuses on the multimodal domain of images and texts, we propose a model that can learn the matching relationship between text-image pairs by contrastive learning.More specifically, We train the model from the labeled data provided by the official organizer, after pre-training, texts are used to reference learned visual concepts enabling visual word sense disambiguation tasks.In addition, the top results our teams got have been released showing the effectiveness of our solution.We release our code at https://github.com/Insaner1004/VWSD.