Robot Confrontation Based on Policy-Space Response Oracles
Mingxi Hu, Siyu Xia, Chenheng Zhang, Xian Guo · 2022 12th International Conference on CYBER Technology in Automation, Control, and Intelligent Systems (CYBER) · 2022
Solving Nash equilibrium solutions in game has always been a challenging problem, especially for the robot confrontation because of the differential constraint and the complex rewards. In this paper, a novel and stable policy-space response oracles(PSRO) method is proposed which integrates$\alpha$-Rank as the meta-strategy solver and covariance matrix evolutionary preference-based policy search(CMA-EPPS) as the oracles solver. Specifically, the robot confrontation task is formulated as a Zero-sum Game problem. Then a novel policy-space response oracles method is developed to solve the Nash equilibrium.$\alpha$-Rank is used to solve the meta-strategy and to improve the efficiency of the oracles solver, CMA-EPPS is adopted. Finally, the proposed method is utilized to solve the robot confrontation and derives a more stable and better solution than iterated best response, fictitious play and PSRO based on$\alpha$-Rank and Best Response.