A deep reinforcement learning approach with a parallel training scheme for Kinova manipulator

Wen-Yu Cheng, Cameron Veit, Zhen Ni, Erik D. Engeberg, Xiangnan Zhong · 2025

Deep reinforcement learning (DRL) has become a popular machine learning (ML) method for training intelligent agents to perform complex tasks. Proximal policy optimization (PPO) has recently emerged as a leading algorithm in DRL, yet its potential for accessible robotic manipulator control has not been extensively studied. To address this gap, we propose a novel parallel training and implementation framework for the Kinova robotic manipulator for trajectory planning of reach-to-grasp tasks. Specifically, we present a testbed built with Unity as the simulation environment and multiple custom Kinova robotic manipulator models as parallel learning agents. Using an adaptive PPO controls algorithm, each agent learns to automatically to reach random goal locations through trial-and-error. The parameters of each agent are aggregated into a combined model which is further refined during training by individual agent reward contributions. Experiments are conducted to determine the optimal number of parallel learning agents by evaluating average reward, successful steps, accuracy, and training time. Our results show that with 16 agents (manipulators) training simultaneously, the accuracy could reach up to 80%, outperforming the single-agent training performance of 30%, while utilizing less than a third of the training time. The improved performance of our framework, combined with its potential for robotic vision integration, can significantly enhance various robotic operations. This includes future military applications, such as autonomous robotic arms in Explosive Ordnance Disposal (EOD) robots, enabling them to operate more effectively in complex and dynamic environments.

Read the paper · More papers on PaperTik