Knowledge revision for reinforcement learning with abstract MDPs

Kyriakos Efthymiadis, Sam Devlin, Daniel Kudenko⋆ · 2014

Reward shaping is a method often used in RL so as to pro-vide domain knowledge to agents and thus improve learn-ing. An unrealistic assumption however is that the provided knowledge is always correct. This assumption can lead to poor performance in terms of total reward and convergence speed in case it is not met. Previous research demonstrated the use of plan-based reward shaping with knowledge re-vision in a single agent scenario where agents showed that they can quickly identify and revise erroneous knowledge and thus benefit from more accurate plans. This method however has no mechanism to deal with non-deterministic scenarios and is thus limited to deterministic domains. In this paper we present a method to provide heuristic knowl-edge via abstract MDPs, coupled with a revision algorithm to manage the cases where the provided domain knowledge is wrong. We show empirically that our method can effi-ciently revise erroneous knowledge even in the cases where the environment is non-deterministic and also removes the need for some of the assumptions present in plan-based re-ward shaping with knowledge revision.

Read the paper · More papers on PaperTik