SlateLLM: Distilling LLM Semantics into Session-Aware Slate Recommendation without Inference Overhead

Aayush Singha Roy, Ηλίας Τράγος, Aonghus Lawlor, Neil Hurley · 2025

Session-based slate recommendation systems curate ranked sets of items in real-time, adapting to evolving user interactions.Balancing relevance, diversity, and novelty remains challenging for reinforcement learning (RL) methods.Recent advances in large language models (LLMs) offer a new possibility to leverage their semantic reasoning capabilities to refine slate composition.In this work, we examine the impact of LLM-driven reasoning on slate generation by integrating LLMs with an RL-based slate recommender and evaluating in terms of accuracy, similarity, diversity, and novelty.We extend the RecSim framework with real-world interaction data and introduce a session-aware evaluation protocol that captures long-term engagement.Our analysis reveals that LLM reasoning enhances subcategory-level diversity while maintaining relevance, leading to increased user engagement.By visualizing category-level shifts in slate composition we uncover systematic patterns in how LLMs refine recommendation diversity.Although direct LLM use during inference may be hampered by computational demands and latency concerns, our experimental results demonstrate that integrating LLM modifications during training enables the model to internalize the nuanced characteristics of LLM reasoning without incurring inference overhead, thereby improving recommendation performance, serving time efficiency, and deployability.

Read the paper · More papers on PaperTik