Model-free safe policy learning via hard action barrier functions
Agustin Castellano, Juan Andrés Bazerque, Enrique Mallada · 2021
To enable model-free learning of safe policies in Constrained Markov Decision Processes, we advance the notion of penalties as an information source that complements rewards and is similarly acquired from experience. Without prior knowledge of the safety constraints, the agent must resort to penalties-that signal infeasibility-in order to learn which actions lead to constraint violations. Accordingly, we define the notion of hard action barrier functions, which incorporate penalties as (0,+infinity) binary information, and gradually reveal implicit state-action pair constraints that must be satisfied in order to avoid (possibly future) unsafe states. Using this notion, we characterize a separation principle that decouples safety from reward maximization. Based on this principle we propose an adaptive algorithm that learns the action barrier function independently of the specific reward structure. As a result, our Barrier-Learning algorithm can wrap around standard on-and off-policy algorithms such as Q-Learning and SARSA. Our solution has the added benefit of learning from previous mistakes by avoiding bumping into the same rock twice, i.e., not taking the same unsafe action that led to a constraint violation in the past. This results in a policy that complies with the constraints almost surely. We demonstrate these combined algorithms in a grid-walk with walls that must be avoided on the way to a target.