Hierarchical Reinforcement Learning in Communication-Mediated Multiagent Coordination
Felix Raoul Fischer, Michael Rovatsos, Gerhard Weiß · 2004
This paper proposes hierarchical reinforcement learning methods for multiagent coordination problems modelled as Markov Decision Processes (MDP). Starting from the observation that communication can aid in predicting others ’ behaviours, we suggest the use of inter-agent messages to mediate between state transitions in the original MDP. Since message exchange has little effect on the MDP (both consequence- and utility-wise), we are able to reduce the problem of learning an optimal policy for the multiagent MDP to learning an optimal communication policy. To solve this problem for realistic domains, we utilise interaction frames as powerful, knowledge-level policy abstractions that can be combined with case-based reasoning techniques. The approach is validated through experiments in a complex application domain which prove that it is capable of heuristically handling significantly larger state and action spaces than exact MDP solution methods. 1.