Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
Moyu Zhang, Yongxiang Tang, Yujun Jin, Jinxin Hu, Yu Zhang · 2025
Traditional recommendation methods model a user's interest in a target item by correlating its embedding with the embeddings of items from the user's interaction history, thereby capturing implicit collaborative filtering signals. Consequently, traditional ID-based methods often encounter data sparsity problems stemming from the sparse nature of ID features. To mitigate this issue, recommendation models incorporate multimodal item information to enhance recommendation accuracy. However, existing multimodal recommendation methods typically rely on early fusion approaches, which focus primarily on combining text and image features, while neglecting the dynamic context provided by user behavior sequences. This oversight precludes the dynamic adaptation of multimodal interest representations to behavioral patterns, thereby hindering the model's ability to effectively capture user multimodal interests. Therefore, this paper proposes the Distribution-Guided Multimodal-Interest Auto-Encoder (DMAE), which achieves the cross fusion of user multimodal interest at the behavioral level. Specifically, DMAE comprises three key components: 1) Multimodal Interest Encoding Unit (MIEU), which encodes the similarity scores between the target item and historically clicked items as the corresponding representation vectors of user interest across different modalities. 2) Multimodal Interest Fusion Unit (MIFU), which dynamically adapts these interest representations through both intra- and inter-modal fusion, a process contextualized by the user's behavioral sequence to achieve a fine-grained and behavior-aware representation of interest. 3) Interest-Distribution Decoding Unit (IDDU), which employs a decoder to reconstruct the encoded user interest representations into true similarity distributions for each modality. The similarity distributions serve as a guide for model learning, aiming to retain as much multimodal information as possible. Ultimately, extensive experiments demonstrate the superiority of DMAE.