Multimodal Intent Recognition in E-Commerce: Challenges, Innovations, and Lightweight Solutions
Junwen Liu, Haikuan Huang, Gang Huang, Shuang Ge, Jinlian Hu · 2025
The "Multimodal Dialogue System Intent Recognition Challenge", jointly organized by Alibaba Taobao and Tmall Group and the World Wide Web Conference (WWW), focuses on multimodal intent recognition in e-commerce scenarios, aiming to address technical challenges in joint understanding of images and text for customer service applications. During the preliminary round, over 1,500 teams participated, with 11 advancing to the semi-finals and 9 ultimately presenting in the final. Innovative approaches from participants centered on three key directions: multimodal data augmentation (e.g., synthetic sample generation, image-text co-augmentation), model optimization (discriminative fine-tuning, model soup fusion, etc.), and prompt engineering. Significantly, the top three teams elevated the weighted F1-score from a baseline of 0.78 to above 0.9. This improvement was achieved by incorporating a diverse set of techniques, including but not limited to vision-language models, structure-aware retrieval, and hierarchical label optimization. The competition outcomes validate the potential of lightweight models in data-scarce scenarios and provide open-source technical pathways for applying multimodal large language models to e-commerce customer service. These advancements drive progress in fine-grained semantic comprehension, domain adaptation, and efficient inference, offering valuable insights for the industrial deployment of intelligent customer service systems.