Multi-modal Knowledge Graph Recommendation based on Self-Supervised Learning

Congying Yan, Qingwei Xia, Yuhan Hu, Li Li · 2024

Incorporating multi-modal information into user-item interaction graphs has emerged as a promising strategy to improve recommendation quality by capturing richer semantic relationships. This paper introduces a novel framework that uses bootstrap latent representations improved by contrastive learning to integrate multi-modal knowledge graphs into the recommendation process. Contrastive learning addresses the cold-start problem by maximizing mutual dependencies between item content and collaborative signals [1]. We construct a comprehensive multi-modal knowledge graph by enriching the user-item interaction graph with structured knowledge, images, and textual data. After encoding both user-item interactions and multi-modal information, we propose a self-supervised learning paradigm that eliminates the need for negative sampling and complex data augmentations, thereby reducing computational overhead. Our model incorporates a contrastive view generator and leverages multiple loss functions—including graph reconstruction loss, inter-modality feature alignment loss, and intra-modality feature masking loss—to learn robust and discriminative representations. Our methodology offers improved accuracy and efficiency compared to existing methods, as demonstrated by experimental results on standard recommendation benchmarks. It also outperforms them significantly. This study highlights the effectiveness of combining multi-modal knowledge graphs and contrastive learning to enhance recommendation systems.

Read the paper · More papers on PaperTik