Policy Interpretation for Deep Reinforcement Learning
Anastasiya Slavova, Vladimir Dimitrov Hristov · 2025
Reinforcement Learning (RL) aims to train an autonomous intelligent agent to build interaction behavior with a predefined environment. This is usually achieved through training using deep neural networks. The Deep Reinforcement Learning (DRL) method is rapidly entering various domains, such as robotics, autonomous cars, industrial production, and other systems implemented in human environments. Due to the black box nature of this training approach, it is important to understand the behavior of the RL agent, especially in areas with risk to human health. Understanding decision making models allows building reliable and stable intelligent systems, find drawbacks, and optimize the learning algorithm. In this work, we develop a method for interpreting RL agent behavior which is based on Agglomerative clustering of Shapley values embeddings of states for gaining insights of agent politics in various resulting groups through RuleFit method. The results obtained point out parameters importance, state thresholds of actions and marginal state values.