Interpretability in Deep Reinforcement Learning

Sarthak Das, Rajarshi Sankar Ray · 2024

Interpretable Machine Learning (IML) has been described as an attempt to understand the behaviour of machine learning algorithms and the rationale behind why a model in question makes a particular prediction for the given input. Interpretability is particularly valuable in Reinforcement Learning (RL) as it is expected to help reduce the RL search space and make RL easier to troubleshoot and use. This work is a literature survey that concerns itself with interpretability techniques directed towards explaining decisions taken by trained Deep RL (DRL) agents in various environments.

Read the paper · More papers on PaperTik