Towards robust shielded reinforcement learning through adaptive constraints and exploration: The fear field framework
Haritz Odriozola-Olalde, Maider Zamalloa, Nestor Arana-Arexolaleiba, Jon Perez-Cerrolaza · Engineering Applications of Artificial Intelligence · 2025
Machine Learning (ML) techniques, including Reinforcement Learning (RL), demonstrate potential as decisionmaking controllers.However, enhancing the robustness required for real-world deployment remains imperative.Within the realm of Safe RL, Shielded RL emerges as a solution, employing shields to block actions leading to unsafe states and offering safe alternatives through known policies.Yet, many Shielded RL methods rely on dynamic environment models, which may inaccurately predict future states, compromising controller robustness.We introduce the Fear Field framework to mitigate this issue for discrete Markov Decision Processbased (MDP) shields with strictly connected unsafe state spaces and fully observable states, which adjusts safe operation constraints based on disparities between model predictions and actual environmental dynamics.We employ parallel learning and Curriculum Learning (CL) strategies to mitigate lengthy training times in high state-space size environments.Additionally, an adaptive exploration algorithm enhances convergence rates amidst significant environmental dynamic shifts.In our case study, integrating CL and the adaptive exploration algorithm with the Fear Field framework reduces unsafe state occurrences by two orders of magnitude while enhancing convergence time following sudden environmental changes.The Fear Field framework significantly reduces unsafe states in the Frozen Lake Gridworld environment at low computational expense when model predictions deviate from reality, with negligible costs otherwise.