PPO-Based Mobile Robot Path Planning with Dense Bootstrap Reward in Partially Observable Environments

Shujie Zhou · 2024

Path planning in large, complex environments poses significant challenges, as a single robot must efficiently navigate to its target while avoiding obstacles. Traditional algorithms often struggle to balance computational efficiency and planning accuracy. Reinforcement learning methods, while promising, frequently suffer from convergence issues due to difficult-to-tune hyperparameters and sparse reward signals. This paper introduces a Proximal Policy Optimization (PPO)-based path planning approach with a dense bootstrap reward mechanism. The proposed method enhances learning efficiency and ensures smooth navigation by incorporating a structured reward design, significantly improving the agent's path-planning performance in complex scenarios. We conducted simulation experiments in a classical warehouse environment to demonstrate the effectiveness and superiority of our proposed method. Compared to traditional rewards, our method significantly improves the pathfinding efficiency of deep reinforcement learning algorithms.

Read the paper · More papers on PaperTik