Don’t Get Bored: Enhancing Scalability and Diversity in Session-Based Slate Recommendation
Aayush Singha Roy, Edoardo D’Amico, Ηλίας Τράγος, Aonghus Lawlor, Neil Hurley · ACM Transactions on Recommender Systems · 2025
Reinforcement learning (RL) has demonstrated great potential to improve slate-based recommender systems by optimizing long-term user engagement. However, addressing the combinatorial action space in slate recommendations remains challenging. Recent work decomposes slate Q -values into item-wise Q -values, improving the tractability of value-based methods to learn the model. But in scenarios with a large item pool and a resource-intensive value function like deep neural networks, the action selection process still incurs substantial computational costs. Slow training might be tolerable, but high costs during action selection could hinder real-time deployment. To address this issue, this article introduces an actor method that reduces Q -function evaluations to a subset of items, significantly cutting inference time for practical deployment. The research suggests acquiring representations at both item and slate levels, strategically identifying a specific item subset for slate composition. The proposed methodologies are assessed over different simulated user engagement behaviors: users certain about preferences (“decisive” behavior) and those more exploratory or bored users, losing interest with repetitive content exposure (“explorative” behavior). Empirical evaluation shows that the proposed approach achieves comparable user engagement with a value-based policy across behaviors. Meanwhile, it notably enhances serving time while recommending diverse topic slates, thus demonstrating its potential effectiveness and efficiency in real-world applications.