Optimized Subgoal Generation in Hierarchical Reinforcement Learning for Coverage Path Planning
Yijun Zhang, Zhiming Li, Ku Du · Automation · 2025
Hierarchical Reinforcement Learning (HRL) for UAV Coverage Path Planning (CPP) is hindered by the “subgoal space explosion”, causing inefficient exploration. To address this, we propose a two-stage framework, Hierarchical Reinforcement Learning Guided by Landmarks (HRGL), which synergistically combines HRL with a multi-scale observation space. The framework provides a low-resolution global map for the high-level policy’s strategic planning and a high-resolution local map for the low-level policy’s execution. To bridge the information gap between these hierarchical views, the first stage, ACHMP, introduces a learned Adjacency Network. This network acts as an efficient proxy for local feasibility by mapping coordinates to an embedding space where distances reflect true reachability, allowing the high-level policy to select feasible subgoals without processing complex local data. The second stage, HRGL, further introduces a landmark-guided global guidance mechanism to overcome local myopia. Extensive experiments on a variety of simulated grid-world maps demonstrate that HRGL significantly outperforms baseline methods in terms of both convergence speed and final coverage rate.