Centralized and Accelerated Multiagent Reinforcement Learning Method with Automatic Reward Setting

Kaoru Sasaki, Hitoshi Iima · Transactions of the Institute of Systems Control and Information Engineers · 2022

For multiagent environments, a centralized reinforcement learner can find optimal policies, but it is time-consuming. A method is proposed for finding the optimal policies acceleratingly, and it uses the centralized learner in combination with supplemental independent learners. In order to prevent the failure of learning, the independent learners must stop in a timely manner, which is done through finely tuning a reward. The reward tuning, however, requires additional time and effort. This paper proposes a reinforcement learning method in which the reward is automatically set.

Read the paper · More papers on PaperTik