Security Reinforcement Learning Guided by Finite Deterministic Automata
Yiming Shi, Qinglong Wang · 2024
In response to the demand for intelligent agents to perform complex tasks in dynamic and uncertain environments. Reinforcement learning algorithms unveil policies that aim to maximize rewards, yet they do not inherently ensure safety throughout the learning or execution phases. To address this issue, we propose a new approach to obtain the optimal strategy while enforcing the temporal logic transformation into the attributes expressed in finite deterministic automata.To achieve this objective, we propose synthesizing a reactive system known as a shield based on the given temporal logic specification that must be adhered to by the learning system. The shield diligently observes the learner's actions and intervenes solely when the selected action deviates from the prescribed specification. Ultimately, we showcase the remarkable adaptability of our approach across various demanding reinforcement learning scenarios.