t-RELOAD: A REinforcement Learning-based Recommendation for Outcome-driven Application

Debanjan Sadhukhan, Sachin Kumar, Swarit Sankule, Tridib Mukherjee · 2023

Games of skill provide an excellent source of entertainment to realize self-esteem, relaxation and social gratification. Engagement in online skill gaming platforms is however heavily dependent on the outcomes and experience (e.g., wins/losses). A user can behave differently under different win/loss experience. An intense engagement can lead to potential demotivation and consequential churn, while a lighter engagement can lead to more confidence for longer sustenance—all depending on the outcomes. Generating a relevant recommendation using reinforcement learning (RL) that can also lead to high engagement (both intensity and duration) is non-trivial because: (i) an early exploration through online-RL using a combined multi-objective reward can permanently hurt users; and (ii) a simulation environment to evaluate RL policies is hard to model due to the (unknown) outcome-driven natural volatility in user behaviour. This work addresses the question “how can we leverage off-policy data for recommendation to solve cold-start problem while ensuring reward-driven optimality from platform-perspective in outcome-based applications? ”. We introduce t-RELOAD: A REinforcement Learning-based REcommendation framework for Outcome-driven Application consisting of 3-layer-based architecture: (i) off-policy data-collection (through already deployed solution), (ii) offline training (using relevancy) and, (iii) online exploration with turbo-reward (t-reward, using engagement). We compare the performance of t-RELOAD with an XGBoost-based recommendation system already in-place to capture the effectiveness.

Read the paper · More papers on PaperTik