Learning, Cooperation, and Coordination in Multi-Agent Systems
H.R. Berenji, David Vengerov · 2000
This paper describes the work on multi-agent learning done at IIS Corp. We first describesingle agent learning algorithms that are used in our multi-agent simulations. We areprimarily interested in dynamic, complex, and uncertain domains where agents do notknow a priori the exact structure of the environment and have to learn it from experience.Agents also do not have a teacher who can tell them which actions are optimal in everysituation that can arise. Therefore, in parallel with learning about the environment, agentsalso have to learn to improve their behavior policy. This problem is naturally formulatedas a reinforcement learning problem (Sutton and Barto, 1998, Bertsekas and Tsitsiklis,1996). In section 3 we describe two reinforcement learning algorithms that can be used inreal world problems with continuous multi-dimensional state spaces and discrete orcontinuous actions spaces.However, individual learning is not sufficient for complex missions of autonomousagents. Even when an agent’s action policy is properly initialized, it can still convergeduring learning to a locally but not globally optimal policy that ignores many importantaspects of the environment. Cooperation among agents during learning is essential indirecting the adjustment of policies in the globally most beneficial direction. In (Berenjiand Vengerov, 2000a), attached in Appendix, we have presented analytical examples ofhow agents learning individually can get stuck with using severely suboptimal policiesand how sharing experience between agents during learning can help to eliminate thisproblem. Our work on cooperative learning in multi-agent systems is described in section4.In addition to cooperative learning among agents working on similar tasks, coordinationof actions in larger teams is essential. When significant uncertainty is present in theenvironment, coordination among agents has to be adaptive. Agents need to dynamicallyallocate responsibilities for different subtasks depending on the context of each agent aswell as changing circumstances of the overall situation. As different critical situationsarise during a remote mission, agents need to be able to handle them with a higherpriority than in the normal mode of operation. Also, as agents arrive at an unexploredarea of the search space, they need to be able to specialize their behavior to therequirements of the situation, delegating more agents to handle more prevalent or moreimportant aspects.In section 5 we describe our work on adaptive coordination in multi-agent reinforcementlearning problems. We consider two related scenarios, which form a natural hierarchy. Inthe first scenario considered in section 5.1, agents learn to distribute themselves in spaceso as to match the dynamically changing distribution of the underlying resources. In thesecond scenario considered in section 5.2, agents can become experts in detecting andreacting to certain classes of situations that can arise in the environment. In order tomaximize the team benefit, agents need to learn to distribute themselves among theclasses of expertise so as to match the dynamically changing importance of each class.