MDN: Modality Decomposition Network for Multimodal Recommendation
Zhuoyang Liu, Weihai Lu · 2025
With the rapid growth of multimedia applications and content, multimodal recommendation systems have garnered significant attention due to their ability to leverage diverse data types for personalized recommendations. Existing methods, which primarily focus on extracting common features across modalities, encounter two critical limitations: (1) they often overlook modality-unique features that carry distinct and valuable information, and (2) they fail to effectively capture cooperative interactions between modalities, which are essential for comprehensive understanding. To address these challenges, we propose the Modality Decomposition Network for Multimodal Recommendation (MDN). MDN introduces a novel Multimedia Knowledge Decomposition module that systematically separates modality representations into three key components: common features, unique features, and cooperative features. This decomposition enables our model to learn richer and more comprehensive representations by explicitly modeling the interplay between shared and modality-unique information. Additionally, MDN incorporates a Multimodal Information Encoder to enhance item feature representation by integrating diverse data sources. Furthermore, a Multimodal Contrastive Enhancement Layer is designed to refine user and item representations through contrastive learning, ensuring robust and discriminative recommendations. Extensive experiments conducted on benchmark datasets demonstrate that MDN consistently outperforms existing state-of-the-art methods, achieving superior performance.