Comparing Reward Policies in Multi-Agent Reinforcement Learning for 2D Tag Game
Sanghun Jeong, Ku‐Jin Kim · Asia-pacific Journal of Convergent Research Interchange · 2025
This paper proposes and compares various reward policies for implementing a 2D tag game using multi-agent deep reinforcement learning (MADRL) techniques.Tag is a game where a chaser chases and tags a runner, with the runner aiming to avoid being tagged for as long as possible.The game can be customized with variations such as setting a tag limit, adjusting the environment with obstacles, or altering the number of players on each team.Designing an effective reward policy for an agent in a tag game to maximize survival time can be challenging and involve a substantial amount of experimentation.This paper investigates the impact of various reward policies on agent survival time and presents a methodology for identifying the optimal reward policy.Reinforcement learning experiments were conducted on four distinct stages, each with a unique obstacle layout, where agents were trained to maximize survival time while being chased under predefined rules.Eleven different reward policies were applied to evaluate their effectiveness, each focused on avoiding collisions with walls, obstacles, and chasers.After determining the optimal reward policy that maximizes agent survival time through experiments, a tag game was implemented where human players chase agents trained under this policy.