Ameliorating Learning Stability by Regularizing Soft Actor-Critic

Junqiu Wang, Xiang Feng, Bo Liu · 2024

Off-policy reinforcement learning algorithms have been successfully applied in different kinds of control tasks such as locomotion for quadrupedal robots and robotic manipulation. Despite of these achievements, deep reinforcement learning algorithms are facing problems including insufficient sample efficiency and instability during training. In this work, we partially address these challenges by considering entropy augmented exploration and adaptive sampling. We extend the ideas in Soft Actor-Critic (SAC) for encouraging exploration. We apply Markov Decision Process regularization in two aspects. First, we penalize deterministic policy setting using information entropy constraints and provide a more reasonable entropy target. Second, we harness an effective sampling method from the replay buffer to reduce the bias in Q-function estimations. The proposed algorithm prefers more stochastic policies in the early stage of the learning process. We perform systematic experiments on a range of benchmark control tasks including challenging robotic manipulation and locomotion. The experimental results outperform SAC in sample efficiency. In addition, the proposed algorithm demonstrates good learning stability with different random seeds. The results suggest that our algorithm is a possible alternative candidate for robotic tasks.

Read the paper · More papers on PaperTik