Model-based reinforcement learning in dynamic environments
Marco Wiering · Utrecht University Repository (Utrecht University) · 2002
We study using reinforcement learning in particular dynamic environ-ments. Our environments can contain many dynamic objects which makes optimal planning hard. One way of using information about all dynamic objects is to expand the state description, but this results in a high di-mensional policy space. Our approach is to instantiate information about dynamic objects in the model of the environment and to replan using model-based reinforcement learning whenever this information changes. Furthermore, our approach can be combined with an a-priori model of the changing parts of the environment, which enables the agent to optimally plan a course of action. Results on a navigation task in a Wumpus-like environment with multiple dynamic hostile spider agents show that our system is able to learn good solutions minimizing the risk of hitting spi-der agents. Further experiments show that the time complexity of the algorithm scales well when more information is instantiated in the model.