Adaptive Natural Policy Gradient in Reinforcement Learning

Dazi Li, Zengyuan Qiao, Tianheng Song, Qibing Jin · 2018 IEEE 7th Data Driven Control and Learning Systems Conference (DDCLS) · 2018

In recent years, the policy gradient method in intensive learning has attracted wide attention with its good convergence performance. At the same time, regulation of hyper parameters is also a matter of concern. Based on the advantages of Actor-Critic structure (AC), the Natural-Gradient Actor-Critic algorithm (NAC) in the discount model is studied in this article. Then the Natural-Gradient Actor-Critic with ADADELTA (A-NAC) algorithm is proposed .The use of ADADELTA is adapted to adjust the learning rate in the actor network, and further improves the convergence speed of the NAC algorithm. Simulation results show that NAC/A-NAC have better learning efficiency and faster convergence rate than regular gradient AC methods.

Read the paper · More papers on PaperTik