Accelerating the Deep Reinforcement Learning with Neural Network Compression
Hongjie Zhang, Zhuocheng He, Jing Li · 2019
Acceleration in deep reinforcement learning has attracted the attention of researchers. Many parallel training frameworks have been proposed to speed up the sampling and the training by running multiple agents simultaneously. However, the bottleneck of performance in a single agent is the neural network prediction which takes a lot of time as compared to the rapid environmental changes. Different from these parallel frameworks, we try to speed up the prediction of the agent to accelerate the entire training process. As far as we know, this is the first time to accelerate the training from this perspective. We propose a novel training framework NNC-DRL to accelerate the whole training process of deep reinforcement learning. NNC-DRL uses different neural networks to represent target policy and behavior policy. The behavior policy network is a smaller neural network derived from the target policy network, which could speed up the prediction. The target policy network will transfer its latest policy to the behavior policy network. The inconsistent policy distribution between behavior network and target network will degrade the convergence of the training. In NNC-DRL, we introduce the Important Sampling technique to estimate policy gradient, which could improve the convergence of the training. The experiments show that our approach NNC-DRL can speed up the whole training process by about 10-20% on Atari 2600 games with little performance loss.