Safe Deep Reinforcement Learning Control with Self-Learned Neural Lyapunov Functions and State Constraints*
Périclès Cocaul, Sylvain Bertrand, Hélène Piet-Lahanier · 2024
In this paper, a Deep Reinforcement Learning (DRL) algorithm is proposed, to learn a control policy for dynamic systems with input and state constraints. The system dynamics are assumed to be unknown for control design, the only available a priori information being the formulation of input and state constraints. Elements of Control Theory are leveraged to obtain safety certificates along with the learned policy: Control Lyapunov Functions (CLF) for closed loop stability, and Control Barrier Functions (CBF) for state constraint satisfaction. These two notions are transformed into conditions to be numerically verified in a PPO-based DRL algorithm, where a neural network CLF is also learned along with the control policy. Simulation results are presented for two application examples, illustrating closed-loop stability and constraint satisfaction resulting from the learned controllers.