Dual Models to Facilitate Learning of Policy Network

Yidong Mei, Yue Gao, Shaoyuan Li · 2020

Learning to control agents without the prior knowledge of its own kinematic is a challenging problem. Recently neural network architecture models can be utilized for robust nonlinear fitting. However, complex sensor inputs including RGB images and robot joints states bring difficulties to the convergence of the control policy due to large input space and complex network structure. Besides, adding temporal information can bring more perception capabilities to the agent but also increase the number of parameters for policy networks. We present a new method called DualM-Control, which exploits dual models including self model and image model. DualM-Control algorithm can compress high dimensional spatial and temporal sensor inputs into low dimension data, thus making it possible for policy network with a few thousand parameters to evolve with evolution strategy. We test the algorithm on a challenging simulation environment created on gym and the performance exceeds existing approaches.

Read the paper · More papers on PaperTik