Unbiased Meta Reinforcement Learning for Interactive Recommender Systems
Huiting Liu, Xinlong Lv, Peng Zhao, Peipei Li, Xindong Wu · IEEE Transactions on Multimedia · 2025
Interactive recommender systems have garnered widespread attention due to their ability to dynamically update recommendation strategies based on user feedback, enhancing the user's interactive experience. To maximize long-term user satisfaction, existing research has incorporated reinforcement learning into interactive recommender systems and combined it with meta-learning to form a meta-reinforcement learning framework that further addresses the cold-start problem in interactive recommendation. However, on one hand, there are latent confounders affecting user feedback; on the other hand, since training samples are observed rather than experimentally obtained, selection bias and exposure bias exist in the interactive data. Most existing studies remove biases using the method of Inverse Propensity Score, which often utilizes fixed propensity scores and neglects the latent confounders affecting user feedback. In this paper, we propose an unbiased interactive recommender system (UIRS) based on a meta-reinforcement learning framework. To eliminate the impact of latent confounders in the state encoding process, we design a user preference representer consisting of three interconnected gated recurrent units. Additionally, we use the item recommendation probabilities output from the policy network as propensity scores and design the objective functions based on these scores, to eliminate biases while addressing latent confounders. Extensive experiments conducted on three benchmark datasets demonstrate that our proposed UIRS model achieves significant improvements over existing state-of-the-art baseline models.