Speeding Up Reinforcement Learning by Exploiting Causality in Reward Sequences
Hongming Li, José Carlos Príncipe · 2021
This paper presents a methodology to exploit causation in deep reinforcement learning (DRL). We take advantage of a cognitive architecture that automatically decomposes the game world in proto-objects and records their position in 2D space across time. Therefore, while playing the game, proto-objects' locations define internal time sequences that can be compared with the reward sequence to select the proto-object that caused the reward. We propose a novel non-parametric information theoretic learning Granger causality (ITL-GC) estimator of directed information using Reny's entropy that is accurate in high dimensions. We integrate this module in a state-of-the-art DRL architecture (A3C) and show substantial improvement in the speed of convergence compared with conventional training.