2A1-L03 The reward distribution based on peripheral information for multi-agent reinforcement learning
Tomoya KIMURA, Yasuo Kuniyoshi · The Proceedings of JSME annual Conference on Robotics and Mechatronics (Robomec) · 2015
In multi-agent reinforcement learning where the whole system receives one reward, it is necessary for effective learning to distribute reward for each agent appropriately. In this paper, we define a value called "confidence factor", and suggest a novel reward distribution method using this value. In this method, each agent calculates own confidence factor for the success of the task using its peripheral information. Then, the contribution factor of each agent is estimated and the reward is distributed between agents. Using the peripheral information of each agent, effective reduction of the observation cost of state-space and appropriate reward distribution becomes possible. By simulation tasks, we empirically show the effectiveness of the reward distribution of proposed method.