Research on Proximal Policy Optimization Algorithm Based on N-step Update

Zhao Guoqing, Jun‐Ming Xu, Liu Ai-dong, Yu Jing · 2021

PPO algorithm is updated in temporal-difference. Although it is more stable than monte-carlo update algorithm, the iterative cost is greatly increased and the convergence effect is difficult to guarantee. To solve the above problems, an algorithm with N-step updating is proposed to improve it, which is called n-PPO. Specifically, the algorithm not only absorbs the characteristics that temporal-difference updating method has comprehensive exploration space and can estimate the value flexibly and quickly, but also takes into account the advantages that monte-carlo updating method has accurate results, less iterations and fast convergence when exploring the complete state sequence. Experimental results show that the proposed method can reduce the volatility and variance of data under the premise of ensuring correct convergence.

Read the paper · More papers on PaperTik