Reinforcement Learning-Based Recommender Systems Enhanced With Graph Neural Networks

Chenxi Fan, Satoshi Fujita · IEEE Access · 2025

Graph Neural Networks (GNNs) have emerged as powerful tools in recommender systems, enabling the modeling of complex user-item interactions by leveraging graph-structured representations. However, conventional GNN-based recommendation models often struggle to adapt to dynamically evolving user preferences and newly introduced items, as they predominantly rely on static supervised learning frameworks. These limitations hinder their ability to provide accurate and personalized recommendations over time, particularly in scenarios where user behavior and item availability change frequently. To address this challenge, we propose a novel recommender system that integrates GNNs with Reinforcement Learning (RL), combining the structural modeling capabilities of GNNs with the sequential decision-making strengths of RL. Our approach enables continuous learning and adaptation to shifting user preferences while optimizing long-term user engagement. Specifically, we introduce three key components. First, a self-attention-based state encoder adaptively captures both long-term and short-term user preferences from historical interactions. Second, a residual connection is added to the actor network to align the action space with the user embedding space, thereby stabilizing training and improving representation consistency. Third, we design a hybrid reward function that combines user feedback types (e.g., clicks, cart, purchase) with embedding-based similarity, enabling the model to balance accuracy, diversity, and adaptability. We evaluate our method on two real-world datasets: Taobao and Amazon-CD, using Recall, Hit Ratio, NDCG, Item Coverage, and Entropy-based Diversity as core evaluation metrics. Experimental results show that our method consistently outperforms baselines: it improves Recall by 21.4% and 50.4%, Hit Ratio by 8.3% and 33.9%, and NDCG by 11.0% and 22.2% on Taobao and Amazon-CD, respectively, while maintaining competitive Coverage and Entropy scores. These results demonstrate the effectiveness and robustness of our approach in dynamic recommendation scenarios.

Read the paper · More papers on PaperTik