An end-to-end learning of driving strategies based on DDPG and imitation learning
Qijie Zou, Kang Xiong, Yingli Hou · 2020
The Deep Deterministic Policy Gradient Algorithm (DDPG) has great advantages in continuous control problems and plays a very important role in the field of autonomous driving. However, the performance of the traditional DDPG algorithm depends on the initialization parameter settings, and it is difficult to achieve satisfactory results in the actual environment. At the same time, traditional DDPG also needs a lot of exploration to converge to a suitable control strategy. This paper proposes a DDPG algorithm framework based on imitation learning (DDPG-IL). The framework first obtains demonstration data (via IL) and stores it in the expert pool, meanwhile the DDPG algorithm is pre-trained. Then the algorithm makes reasonable use of the demonstration data and its own exploration data for learning. Finally, when the algorithm reaches an approximate expert level, it gradually becomes an ordinary reinforcement learning and continues training by self-learning until the algorithm converges to a stable state. The experimental comparison on the racing simulator TORCS proves that our proposed DDPG-IL algorithm has more advantages and can obtain better performance than the traditional DDPG algorithm.