Hindsight Reward Shaping in Deep Reinforcement Learning

Byron de Villiers, Deon Sabatta · 2020 International SAUPEC/RobMech/PRASA Conference · 2020

Recent developments in the field of deep reinforcement learning (DRL) have shown that reinforcement learning (RL) techniques are able to solve highly complex problems by learning an optimal policy for autonomous control tasks. Although RL shows great promise in sequential decision-making problems in dynamic environments, there are still caveats associated with the framework. One such pitfall is the time taken to converge due to sparse and delayed rewards, known as the temporal credit assignment problem in RL. This paper addresses the problem by introducing a simple yet effective method of distributing the discounted terminal state reward backwards in time for episodic environments after an episode has reached terminal state. The shaped reward transitions are then added to the experience replay buffer.

Read the paper · More papers on PaperTik