Multi-Agent Cooperation by Q-Learning in Continuous Action Domain
Kao‐Shing Hwang, Yu-Hong Lin, Chia-Yue Lo · 2008
In this paper we propose Q-learning with continuous action space and extend this algorithm to a multi-agent system. Conventional Q-learning needs a pre-defined and discrete state space. But it is not practical because the states of the environment in the real world and actions are both continuous. The algorithm will use a concept that is similar to the SRV (Stochastic Real-Valued Unit) to train the actions in each state. The convergence of the SRV may fall into local solution even if it has never reached the optimal solution. In order to overcome this drawback, the Q-learning with SRRV (Stochastic Recording Real-Valued unit) is proposed, and it shows that the SRRV will converge more quickly.