Research on Manipulator Control Based on Improved Proximal Policy Optimization Algorithm
Shaoxiong Yang, Di Wu, Yan Pan, Yan Xiang He · 2022 34th Chinese Control and Decision Conference (CCDC) · 2022
In the scene of random patching in the industrial scene, an algorithm based on a distributed frame of proximal policy optimization (PPO) with Generalized Advantage Estimation (GAE) is proposed in this paper. The visual part is taken from camera, which is considered as state input. A distributed approach (actor-critic) is established to improve the efficiency of sampling. The sampling data are stored in the experience pool. Both punishment and reward strategies are considered in the raised method. The improved PPO algorithm can be verified on Pybullet. We found that it greatly improves effect in terms of convergence steps and actual reward performance.