Reinforcement learning with constraint based on mirror descent algorithm
Megumi Miyashita, Toshiyuki Kondo, Shiro Yano · Results in Control and Optimization · 2021
An important issue in reinforcement learning is to make the agent avoid the dangers and risks during the task such as physical collisions. We propose the reinforcement learning algorithm based on the CoMirror algorithm, named CoMDS, for the problem that has a functional constraint. Besides, we modify the proposed algorithm CoMDS to Gaussian CoMDS for practical use. We evaluate our algorithms with the via-point task of a planar robotic arm with a forbidden area, that employs as a constraint, in the simulator. As a result, we find that Gaussian CoMDS explores the policy while satisfying the constraint.