Mapless Navigation Based on VDAS-PPO Deep Reinforcement Learning

Jiajun Wu, Weihao Chen, Jiaming Ji, Xing Chen, Lumei Su, Houde Dai · 2022

Low sample utilization is a problem for the continuous action space’s Proximal Policy Optimization (PPO) algorithm. This paper proposes a Proximal Policy Optimization (VDAS-PPO) based on Velocity-Directed Action Selection. Firstly, the multi-information fusion reinforcement learning network is designed for mapless navigation to well utilize sample experience and information representation in the current state by fusing target location information, velocity value information, and lidar observation value information. Moreover, during the decision-making process, a novel reward function is proposed to optimize the action decision by designing a speed variation factor. The sample utilization rate is significantly increased by the VDAS-PPO algorithm, which also guarantees that the agent can learn quickly and effectively when training the network model. Three selected scenarios are built up on the simulation platform Gazebo to test the effectiveness of the proposed VDAS-PPO algorithm. The training episodes of the algorithm in this paper are decreased by 5–6 times compared to the PPO and Deep Q Network (DQN) algorithms, and the SPL value is also enhanced.

Read the paper · More papers on PaperTik