Multiagent Simulation On Hide and Seek Games Using Policy Gradient Trust Region Policy Optimization

Hani'ah Wafa, Judhi Santoso · 2020

The hide and seek game is a game that implements a multi-agent system so that it will be solved by using multi-agent reinforcement learning. In this research, we examine how to apply policy gradient method, Trust Region Policy Optimization (TRPO) to solve the hide and seek game environment. We also examine the configuration of TRPO algorithm that gives the best performance and its comparison with the Vanilla Policy Gradient (VPG) algorithm. From the results of experiments conducted, we found that in general the VPG algorithm gives a better performance than the TRPO algorithm when tested in the same environment as the training environment. However, on the contrary, the TRPO algorithm gives better performance when tested in a different environment than that in the training environment.

Read the paper · More papers on PaperTik