Distilling Gaming Strategy through Explainability in Tetris
Lagiokapa Eleftheria, Loupas Georgios, Makrina Viola Kosti, Nefeli Georgakopoulou, Sotiris Diplaris, Stefanos Vrochidis · 2025
As artificial intelligence (AI) systems become increasingly integrated into game design, the demand for transparent and adaptive decision-making grows. While Explainable AI (XAI) has illuminated the internal reasoning of AI agents, most explanation-based training methods have traditionally prioritized alignment with a teacher model over the exploration of strategic diversity. In this paper, we introduce a novel framework that leverages explanation-based knowledge distillation to modulate agents’ internal reasoning, yielding both convergent and divergent behavioral strategies. To demonstrate this approach, we conducted experiments in a Tetris environment comparing baseline agents trained with standard reinforcement learning to agents whose training was modified by incorporating explainability losses. Our dynamic framework integrates a feedback mechanism that adjusts the influence of the explainability term based on performance and strategic utility. This work demonstrates the potential of employing explainability not only as an interpretative tool but also as a means to actively diversify and refine strategies in complex, dynamic environments.