M3Rec: Cross-Modal Context Enhanced Micro-Video Recommendation with Mutual Information Maximization

Qihan Du, Li Yu, Huiyuan Li, Ningrui Ou, Xinjing Gong, Junyao Xiang · 2022 IEEE International Conference on Multimedia and Expo (ICME) · 2022

Personalized recommendation of micro-videos is crucial for many content sharing platforms. Micro-videos typically com-prise fruitful multimodal contents. Existing methods have limitations in learning multimodal representations: (i) They suffer from data sparsity problems as only rely on the inter-action prediction loss to train the whole model. (ii) They do not consider the correlation among multimodal contents into the representation learning. To tackle these limitations, we propose a cross-Modal context enhanced method via Mutual infoMax for Recommendation, termed M3Rec. Specifically, we design the cross-modal graph neural network to inter-change multimodal information to generate modality-aware representations. Then, based on the coupled modality con-text, we design the masked modality prediction (MMP) with three self-supervised objectives to learn the correlation among visual, textual, and acoustic contents empowered by the mu-tual information maximization principle. Finally, we enhance representations via self-supervised pre-training to boost rec-ommendation. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art methods.

Read the paper · More papers on PaperTik