Off-policy Learning over Heterogeneous Information for Recommendation

Xiangmeng Wang, Qian Li, Dianer Yu, Guandong Xu · Proceedings of the ACM Web Conference 2022 · 2022

Reinforcement learning has recently become an active topic in recommender system research, where the logged data that records interactions between items and users feedback is used to discover the policy. Much off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has been a popular research topic in reinforcement learning. However, the log entries are biased in that the logs over-represent actions favored by the recommender system, as the user feedback contains only partial information limited to the particular items exposed to the user. As a result, the policy learned from such off-line logged data tends to be biased from the true behaviour policy.

Read the paper · More papers on PaperTik