EcomMIR: Towards Intelligent Multimodal Intent Recognition in E-Commerce Dialogue Systems

Tianhong Gao, Genhang Shen, Yuxuan Wu, Zunlei Feng, Jinshan Zhang, Sheng Zhou · 2025

Image scene classification and dialogue intent recognition are fundamental tasks in intelligent e-commerce. The former classifies user-uploaded images, while the latter integrates multi-turn dialogue and visual information to extract user intent. However, existing multimodal models struggle with effectively utilizing e-commerce data and generalizing to vertical domains, limiting their practical applicability. To address this, we propose EcomMIR, an E-COMmerce Multimodal Intent Recognition framework based on CN-CLIP and MiniCPM-V. EcomMIR improves model generalization and robustness through multi-level intent data denoising, high-confidence data selection, and hierarchical labeling. Specifically, CN-CLIP employs contrastive learning to align image and text embeddings for efficient scene classification, while MiniCPM-V, a multimodal large language model, deeply integrates textual and visual information to model dialogue context and accurately recognize user intent. Experimental results show that EcomMIR achieves superior performance in both tasks, ranking Top 2 in the WWW2025 Multimodal Intent Recognition for Dialogue Systems challenge, offering an effective solution for multimodal tasks in intelligent e-commerce.

Read the paper · More papers on PaperTik