Learn Once Plan Arbitrarily (LOPA): Dynamic Observation-Based Deep Reinforcement Learning Method for Global Path Planning in Mountainous Terrain Environment
Shuqiao Huang, Mingxin Hou, Xiaofang Yuan, Xiru Wu, Yaonan Wang, Guoming Huang · IEEE Transactions on Artificial Intelligence · 2025
Deep reinforcement learning (DRL) methods have recently shown promise in path planning tasks. However, when dealing with global planning tasks in mountainous terrain (2.5D) environment, these methods face serious challenges such as poor convergence and generalization. To this end, we propose LOPA (Learn Once Plan Arbitrarily), an enhanced DRL method that learns on a single map yet generalizes to topographically similar terrains. Consequently, it enables path planning across multiple mountainous terrain maps while balancing path distance and energy consumption. Firstly, we analyze the reasons of convergence and generalization problems from the perspective of DRL’s observation, revealing that the conventional design causes DRL to be interfered by irrelevant map information. Secondly, we develop the LOPA which utilizes a novel dynamic observation mechanism to attain an improved capability in focusing on key information of the observation. Such a mechanism is realized by two steps: (1) a dynamic observation model is built to transform the DRL’s observation into two dynamic views: local and global, significantly guiding the LOPA to focus on the key information of the given maps; (2) a dual-channel network is constructed to process these two views and integrate them to attain an improved reasoning capability. Meanwhile, through Rademacher Complexity analysis, we provide theoretical justification for LOPA’s improved generalization capability, demonstrating a lower upper bound on the generalization error. The LOPA is validated through multi-objective global path planning experiments conducted on both simulated and real maps. The results suggest that LOPA has improved convergence and generalization performance, as well as great planning efficiency.