Formation Hunting and Trajectory Optimization of Multi-AUV System Based on Adaptive Multi-Agent Reinforcement Learning
Abdul Bari Butt, Min Li, Rao Atif, Muhammad Haider Abbas · 2025
In this study, we investigate the challenges of formation-based hunting and trajectory optimization for multiple autonomous underwater vehicles (AUVs) in a complex underwater environment. Traditional virtual structure algorithms and leader-follower models often encounter challenges when adapting to dynamic underwater environments and are prone to centralized failures. To overcome these limitations, this study proposes a multi-agent deep deterministic policy gradient (MADDPG) framework with continuous state-action spaces to enhance coordination and decision-making. The proposed method aims to improve the success rate and reduce the completion time of formation-based activities. A complex reward function module is incorporated in the underwater multi-AUV simulation environment. The module addresses key issues critical to efficient underwater search, including navigation, formation maintenance, search efficiency, boundary constraints, and collision avoidance. The effectiveness of the proposed technique is demonstrated by comparing the artificial potential field technique and the deep reinforcement learning (RL) algorithm in the simulation environment. The efficiency of task execution is improved by about 16%, and attains a near-optimal success rate.