Improving Zero-Shot Coordination with Diversely Rewarded Partner Agents

Peilin Wu, Zhenhua Yang, Peng Yang · 2024

Zero-shot coordination studies the training of well-generalizing human-AI coordination agents in the scenario where human data is unavailable. To obtain a coordination agent generalize to unseen humans, prevailing methods generate a population of partner agents as proxy models of human partners and then train a coordination agent with these partner agents. Constructed partner agents are expected to be as diverse as possible to cover a wide range of human behaviors, preventing a distribution shift between training and testing stages. Recent works concentrate on studying effective methods of creating a group of high-reward while diverse partner agents to model unseen human partners. However, the resulting high-reward partner agents do not accurately reflect real-world situations, considering that human decisions are not always optimal and may sometimes even hinder the progression of coordination. Therefore, these studies still struggle to capture the potential characteristics of human partners. In this work, reinforcement learning (RL) and supervised learning (SL) are integrated to train a reward-conditioned policy. By conditioned on different desired rewards, a reward-conditioned policy simulates both low-reward and high-reward partners. Additionally, a reward-bucketed replay buffer and curriculum learning are applied to enhance reward diversity and boost the training of coordination agents. Experiments demonstrate that the proposed reward-conditioned policy is capable of generating agents with different rewards. Moreover, the zero-shot coordination performance of agents trained with these partners surpasses previous methods in the majority of scenarios within the Overcooked human-AI coordination benchmark.

Read the paper · More papers on PaperTik