Safe Reinforcement Learning-Based Dynamic Obstacle Avoidance for Robotic Manipulators
Yangxin Zhao, Jianliang Mao, Chuanlin Zhang · 2025
In many tasks, especially in safety-critical scenarios, ensuring safety is of paramount importance. Simulators offer key advantages by enabling safe exploration, which is essential when training Reinforcement learning (RL) systems in phys-ical environments like human-robot interaction. However, RL algorithms still struggle to succeed beyond simulation, mainly due to insufficient safety guarantees during learning, often leading to failures before the policy reaches its optimal form. To overcome this challenge, we introduce a architecture that in-tegrates temporal modeling with safety constraints. The proposed framework combines: (1) a controller based on model-free RL; (2) a policy network incorporating Long Short-Term Memory (LSTM); and (3) a safety-critical controller designed with Control Barrier Functions (CBF). Our general framework exploits RL to achieve high-performance control without relying on system models, while the LSTM network more effectively captures the temporal dynamics of the environment. In addition, the CBF-based controller not only ensures safety but also constrains policy exploration to enhance learning stability.