Effective Learning Approach for Planning and Scheduling in Multi-agent Domain

Sachiyo Arai, Katia P. Sycara · The MIT Press eBooks · 2000

The point we want to make in this paper is that Profit-sharing; a reinforcement learning approach is very appropriate to realize the adaptive behaviors in a multi-agent environment. We discuss the effectiveness of Profit-sharing theoretically and empirically within a Pursuit Game where there exist multiple preys and multiple hunters. In our context of this problem, hunters need to coordinate adaptively one another to capture all the preys, without sharing information, predefined organization and any prior knowledge around their environment. Pursuit Game itself is very simple but can be extended to a real problem. Our approach, Profit-sharing, is contrastive to other reinforcement learning approaches which are based on Dynamic Programming , suchas Temporal Difference method and Q-learning, in that Profit-sharing guarantees convergence to a effective policy even in domains that do not obey the Markov property, if a task is episodic and a credit is assigned in an appropriate manner. Profit-sharing is also different from Q(1) and Sarsa(1) methods in that it does not need eligibility trace to manage the delayed reward. Though our monolithic implementation here seems to be an impractical in a real word, weneed to discuss the validity of algorithm as a multiagent reinforcement learning context before introducing some structured frameworks into the monolithic method to extend its application. The contribution of this paper is that we introduce Profit-sharing as the effective algorithm in the multiagent domain and report its advantages and limitations without hierarchies.

Read the paper · More papers on PaperTik