Safe Reinforcement Learning Based on Off-Policy Approach for Nonlinear Discrete-Time Systems

Mayank Shekhar JHA, Bahare Kiumarsi, Didier Theilliol · 2024

This paper presents a control barrier function-based method for learning safe optimal controllers for discrete-time (DT) nonlinear systems such that safety, stability, and performance are guaranteed on an infinite time horizon. The paper investigates the fusion of reinforcement learning (RL) and control barrier functions (CBFs) and remains novel in that the approach is developed for DT nonlinear systems and develops off-policy safe RL approach for DT systems. Formulation of a novel generalised safety-aware Hamilton Jacobi Bellman (G-SHJB) equation is proposed to assure safety, stability and optimality during exploitation phase. The$\mathrm{G}$. SHJB is solved in an iterative sense to obtain improved control policy. The invariance of safe set as well as stability and optimality of the system under learnt control law is established using mathematically rigorous proofs and studied using simulation.

Read the paper · More papers on PaperTik