Deep reinforcement learning collision avoidance using policy gradient optimisation and Q-learning

Bishoy H. Mikhail, Shady Ahmed Maged · International Journal of Computational Vision and Robotics · 2020

Usage of trust region policy optimisation (TRPO) and proximal policy optimisation (PPO) 'children of policy gradient optimisation method' and deep Q-learning network (DQN) in Lidar-based differential robots are proposed using Turtlebot and OpenAI's baselines optimisation methods. The simulation results proved that the three algorithms are ideal for obstacle avoidance and robot navigation with the utter advantage for TRPO and PPO in complex environments. The used policies can be used in a fully decentralised manner as the learned policy is not constrained by any robot parameters or communication protocols.

Read the paper · More papers on PaperTik