Mountain UGV path planning via optimized dueling double DQN (D3QN): Structural optimization, path-guided rewards, and phased action policy

Gengchen Liu, Song Gao, Junheng Jiang, Zhangmin Luo, Gang Jiang · Control Engineering Practice · 2026

• We propose a structurally optimized D3QN network, which is proposed to dynamize the batch size with layer normalization for the feature extraction layer, and a hybrid network update mechanism is also introduced. • We propose a reward function that combines Bessel hierarchical A* path guidance and APF method, so that the UGVs will no longer search aimlessly in the space species, but quickly approach the existing better path and gradually optimize it, as well as have good obstacle avoidance ability. • A chaotic annealing multi-phased action selection policy is proposed that divides the whole training into two stages, exploration and exploitation, which makes its exploration faster while the stability of exploitation is guaranteed. • The advantages of the proposed method is preliminarily demonstrated by generating a real 3D mountain scene based on digital elevation model (DEM) and grayscale algorithm. Accurate path planning is particularly important for unmanned vehicles in complex mountainous environments. Compared with two-dimensional terrain, mountainous three-dimensional terrain not only introduces more uncertainty but also interference from dynamic obstacles, which dramatically increases the difficulty of path planning. As such, conventional planning methods often struggle to identify efficient solutions. Although path planning techniques utilizing deep reinforcement learning have provided new strategies for solving such problems, existing algorithms face a variety of challenges, including poor network stability, susceptibility to gradient explosion, insufficient reward guidance, and an imbalance between exploration and utilization. To overcome these issues, this paper introduces three novel contributions. First, the dueling double DQN is structurally optimized, and various techniques are introduced to prevent instability and gradient explosion. Second, a new reward function is developed to combine the Bessel hierarchical A* path guidance algorithm with the artificial potential field method, enabling unmanned vehicles to identify the optimal path while dynamically avoiding obstacles. Finally, a chaotic annealing multi-phased strategy is proposed as an action selection policy, which gradually transitions from the exploration stage to the exploitation stage by optimizing the balance between the two as the learning process advances. In addition, a 3D terrain model based on a real mountain environment was generated using the grayscale map algorithm. A series of simulation experiments were conducted to evaluate the performance of the proposed method, as measured by search efficiency, success rate, and path quality. A comparative analysis and comparison with existing DRL path planning algorithms was also performed to provide additional insights.

Read the paper · More papers on PaperTik