Dynamic potential-based reward shaping

Sam Devlin, Daniel Kudenko⋆ · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2012

Potential-based reward shaping can signicantly improve the time needed to learn an optimal policy and, in multi- agent systems, the performance of the nal joint-policy. It has been proven to not alter the optimal policy of an agent learning alone or the Nash equilibria of multiple agents learn- ing together. However, a limitation of existing proofs is the assumption that the potential of a state does not change dynamically during the learning. This assumption often is broken, espe- cially if the reward-shaping function is generated automati- cally. In this paper we prove and demonstrate a method of ex- tending potential-based reward shaping to allow dynamic shaping and maintain the guarantees of policy invariance in the single-agent case and consistent Nash equilibria in the multi-agent case.

Read the paper · More papers on PaperTik