Increasing sample efficiency in deep reinforcement learning using generative environment modelling

Per‐Arne Andersen, Morten Goodwin, Ole‐Christoffer Granmo · Expert Systems · 2020

Abstract Reinforcement learning is a broad scheme of learning algorithms that, in recent times, has shown astonishing performance in controlling agents in environments presented as Markov decision processes. There are several unsolved problems in current state‐of‐the‐art that causes algorithms to learn suboptimal policies, or even diverge and collapse completely. Parts of the solution to address these issues may be related to short‐ and long‐term planning, memory management and exploration for reinforcement learning algorithms. Games are frequently used to benchmark reinforcement learning algorithms as they provide a flexible, reproducible and easy to control environments. Regardless, few games feature the ability to perceive how the algorithm performs exploration, memorization and planning. This article presents The Dreaming Variational Autoencoder with Stochastic Weight Averaging and Generative Adversarial Networks (DVAE‐SWAGAN), a neural network‐based generative modelling architecture for exploration in environments with sparse feedback. We present deep maze, a novel and flexible maze game‐engine that challenges DVAE‐SWAGAN in partial and fully observable state‐spaces, long‐horizon tasks and deterministic and stochastic problems. We show results between different variants of the algorithm and encourage future study in reinforcement learning driven by generative exploration.

Read the paper · More papers on PaperTik