Unmanned ground vehicle exploration position optimization method based on an improved proximal policy optimization algorithm
Bingkun Wang, Yue Wang, Kai Zhao, Wencheng Wei, Peiqi Kang · IET conference proceedings. · 2025
In complex environments, pre-defined exploration positions of unmanned ground vehicles (UGV) often fail to balance exploration success rate and speed, and lack adaptability to dynamic conditions. To address this challenge, this paper proposes an improved Proximal Policy Optimization (IPPO) algorithm that integrates entropy regularization and a decreasing clipping mechanism to optimize exploration position selection for UGV, enabling autonomous optimal decision-making. First, an exploration environment model is constructed, incorporating both an exploration success rate model and an exploration speed model to quantitatively define states, actions, and rewards for reinforcement learning. Then, entropy regularization is introduced into Proximal Policy Optimization (PPO) to enhance exploration capability and prevent premature convergence, while a decreasing clipping mechanism is adopted to gradually constrain policy updates, improving training stability and convergence accuracy. Finally, simulation experiments are conducted to validate the effectiveness of the proposed method. Results demonstrate that the IPPO achieves faster reward convergence, more stable policy learning, and more accurate value estimation compared with baseline approaches, effectively balancing exploration success rate and speed in uncertain environments, and enabling autonomous optimal exploration position selection for UGV.