Partially Observable Multi-Agent RL with Enhanced Deep Distributed Recurrent Q-Network
Longtao Fan, Yuan-yuan Liu, Sen Zhang · 2018
Many real-world problems are naturally modeled as multi-agent problems, and most of them are partially observable. But multi-agent problems will make the environment became nonstationary from the point of view each agent, which causes the combination of experience replay with IQL to be problematic. DDRQN as a method to solve partially observable RL problems, it didn't solve that. In this article we propose a method based on importance sampling and address DDRQN's disadvantages which can't enable memory replay. Results on the SC2LE environment confirm that this method significantly improve performance compared to original DDRQN