Improve Convergence Speed of Multi‐Agent Q‐Learning for Cooperative Task Planning

Arup Kumar Sadhu, Amit Konar · 2020

This chapter aims at extending traditional multi-agent Q-learning (MAQL) algorithms to improve their speed of convergence by incorporating two interesting properties, concerning exploration of the team-goal and selection of joint action at a given joint state. It begins by reviewing the preliminaries of reinforcement leaning (RL). The motivation of RL is to derive the optimal action at a given environmental state for which the agent would be able to derive the maximum reward. Such formulation of deriving optimal action at a given state based on the learned experience of interaction with the environment has plenty of interesting applications, including generating moves in a game, and complex task-planning and motion-planning of a mobile robot in a constrained environment. The chapter then introduces the proposed fast cooperative multi-agent Q-learning algorithms. The chapter deals with multi-agent cooperative planning algorithms and includes experiments and results.

Read the paper · More papers on PaperTik