Limitations of Simultaneous Multiagent Learning in Nonstationary Environments
Itsuki Noda · 2013 IEEE/WIC/ACM International Joint Conferences on Web Intelligence (WI) and Intelligent Agent Technologies (IAT) · 2013
The relationship between the exploration ratio and achievement of learning under multiagent learning (MAL) conditions in nonstationary environments is investigated in this paper. In MAL, exploration of one agent affects other agents' learning, acting in a manner similar to a noise factor, however, exploration is necessary to acquire suitable behaviors and to catch-up changes in a nonstationary environment. The MAL process is formalized from the viewpoint of the learning of probability distribution, where the purpose of learning is defined as maximization of the probability to select the right choice that provides greater benefit than other choices. For the learning, an agent needs to explore all actions to check and confirm that the right action is taken, especially in a nonstationary environment, which by its natuer may cause a change the right action over time even if other agents do not change their policies. On the basis of the proposed formalization, a simple case of resource sharing problems is investigated to show the existence of learning performance boundaries that limit MAL convergence to the right policy during learning.