Variational Invariant Representation Learning for Multimodal Recommendation

Wei Hong Yang, Haoran Zhang, Li Zhang · Society for Industrial and Applied Mathematics eBooks · 2024

Multimodal recommendation systems are widely used in e-commerce and short video platforms. Compared with basic item attribute information, users are more likely to be attracted by item images and make purchases. Many researchers are beginning to explore the efficient use of multimodal information. Many studies have introduced multimodal information into the model as an auxiliary feature and achieved certain results. However, the extraction of critical information from multimodal data containing complex information needs to be optimized. The problem of false relation learning in multimodal recommendation has not been effectively solved. Therefore, we propose a Variational Invariant Representation Learning (VIRL) method for multimodal recommendation. Specifically, we first propose a variational autoencoder based invariant representation learning module. Based on the joint probability distribution of invariant and variant representations, we use variational autoencoders to learn implicit representations. We then design weakly self-supervised contrastive learning to optimize the invariant representations. Further, we introduce adversarial learning to achieve cross-modal invariant information alignment. We conducted a comprehensive experiment on three real-world data sets, and the experimental results show that our model performs best. Further ablation experiments confirm the validity of variational representation.

Read the paper · More papers on PaperTik