Reinforcement Learning Based Recommendation System: An In-Depth Review of Models and their Limitations

Dhaval Mehta, Nipun Dahiya, Ishaan Atre, Soni Sweta · 2025

This paper surveys the landscape of the Al model approaches and algorithms that have been in use in recommendation systems, which have become an essential requirement in enriched user experiences, ranging from e-commerce to content streaming and social media, among other areas. We focus on a few remarkable techniques: Multi-Armed Bandit, Multi-Agent Reinforcement Learning with Deep Deterministic Policy Gradient (DDPG), Deep Q-Network (DQN), Proximal Policy Optimization, and Twin Delayed DDPG. Every model is discussed in detail for the purpose of demarcating the merits and demerits, besides shedding light on some major findings of past literature. We enumerate some important research gaps in the domain that relate to scalability, adaptability to dynamic user preferences, and context-awareness. We aim to highlight the potential and limitations of these algorithms across various applications, proposing directions for future research to improve recommendation systems.

Read the paper · More papers on PaperTik