Optimizing Discriminative Vision-Language Models for Efficient Multimodal Intent Recognition
Jinwang Song, Zhongtian Hua, Hongying Zan, Yingjie Han, Min Peng · 2025
The WWW2025 Multimodal Dialogue System Intent Recognition Challenge aims to enhance the understanding capabilities of e-commerce intelligent customer service systems, requiring participants to develop intent recognition models that jointly analyze textual and visual information for more precise user intent interpretation. This paper presents the solution proposed by Team Adventure, which proposes an optimized framework based on Vision-Language Models (VLMs), incorporating data augmentation, generative and discriminative fine-tuning, and inference optimization. Experimental results demonstrate that discriminative fine-tuning of VLMs offers significant advantages in training efficiency and inference speed. Furthermore, we introduce the Label-Token Initialization (LTI) method, which effectively mitigates the mismatch between the pre-training and discriminative fine-tuning stages.