Autonomous Underwater Vehicle Safety Path Planning Method Based on Constraint Reinforcement Learning
Hao Lu, Zhuo Wang, Guiqiang Bai, Hongde Qin, Wucan Yang · IEEE Transactions on Vehicular Technology · 2025
The high cost of the Autonomous Underwater Vehicle (AUV) and the complexity and unpredictability of underwater environments highlight safety as a primary concern in AUV path planning. The Leash Actor-Critic (LAC) method based on Constraint Reinforcement Learning (CRL) is proposed in this paper to enhance the safety of decision-making within AUV path planning. First, the actions of the AUV are derived from the CRL framework corresponding to the current state. Subsequently, the safety of the AUV's state and actions is estimated by the Emergency Safe Critic (ESC). Finally, the Lagrange multiplier method is used to combine the ESC with the Actor-Critic framework to optimize the safety of AUV strategies. The LAC method applies the ESC to curtail perilous actions, thereby guaranteeing the safety of the AUV's decision-making process. Enhancements to the Lagrange multiplier method and ESC help to minimize the exploration time for AUVs in high-risk underwater environments. Additionally, the adaptability of the AUV to the environment is improved through a sample-based replay buffer. This method was tested in a hardware-in-the-loop simulation system constructed with reference to ocean current data and terrain data from the South China Sea and the Bohai Sea. The experimental results verified its fast convergence speed, short path length, short travel time, and strong generalization ability in unknown and unstructured underwater environments.