Advancing Explainability in Deep Reinforcement Learning: Novel Frameworks for Transparency, Fairness, and Robustness in Autonomous Decision Making Systems
Akanksh Reddy Muddam, Lavanya Reddy Satti · Journal of Emerging Technologies and Innovative Research · 2025
DRL has been shown to be incredibly successful in autonomous decision-making in many complex fields. Nonethe less, due to the black-box character of DRL models, they are not transparent, fair, or trustworthy, and their use is not possible in high-stakes scenarios. This paper suggests innovative frameworks that develop the concept of explainability in DRL through incorporating global interpretability methods, fairness conscious policy analysis, and the improvement of robustness. We present SILVER, a Shapley value-based interpretable policy framework that is a middle ground between explainability and interpretability. The effectiveness of the latter methods is defined by the fact that the experiments conducted on the benchmark control tasks and autonomous systems demonstrate a high level of transparency, fairness indicators, and stability without negatively affecting their functionality. We make progress towards imple menting self-reliant, trustful decision-making systems to enable us to safely deploy in adverse real-life situations.