An environment model in multi-agent reinforcement learning with decentralized training
Rafał Niedziółka-Domański, Jarosław Bylina · Annals of Computer Science and Information Systems · 2024
In multi-agent reinforcement learning scenarios, independent learning, where agents learn independently based on their observations, is often preferred for its scalability and simplicity compared to centralized training.However, it faces significant challenges due to the non-stationary nature of the environment from each agent's perspective.We investigate if incorporating an environment model in multi-agent reinforcement learning with decentralized training can alleviate the non-stationarity effects caused by the adaptive behaviors of other agents.To do this, we design and implement a custom model-based algorithm and compare its performance with the well-known model-free algorithm (Deep Q-Network).Our algorithm uses an environment model to plan and select actions.However, we do not require the model to be perfect for action selection, allowing it to be learned and improved during training.Our results suggest that integrating environment models into MARL offers a viable solution to mitigate non-stationarity.