A Reward-Informed Semi-Personalized Bandit Approach for Enhancing Accuracy and Serendipity in Online Slate Recommendations

Lukas De Kerpel, Dries F. Benoit · ACM Transactions on Recommender Systems · 2025

Contextual bandits provide a principled framework for personalization in online recommendation settings. However, as these methods tailor recommendation slates to an individual user, they tend to induce overspecialization, yielding homogeneous recommendation lists that limit exposure to diverse content and contribute to more systemic issues such as filter bubbles and echo chambers. To mitigate these effects, recommender systems must complement predictive accuracy with serendipity, providing recommendations that are novel and unexpected while remaining contextually relevant. This study proposes a semi-personalized bandit that, for each item, learns a decision tree to segment users by contextual features and reward patterns, and runs a unique Thompson Sampling policy for each user segment to create recommendation slates. By pooling information across behaviorally similar users and conducting the exploration mechanism at the user segment level, the framework mitigates overspecialization issues and promotes serendipitous recommendations. Moreover, the approach is inherently interpretable, with decision trees revealing decision pathways that define user segments, offering insights into recommendation logic. Experiments across three different online domains show that the semi-personalized framework reduces average regret relative to personalized baselines while improving serendipity in sparse interaction settings. These findings underscore the potential of semi-personalized bandits to improve recommendation quality in complex environments.

Read the paper · More papers on PaperTik