Towards efficient and scalable behaviour and communication learning using multi-agent reinforcement learning
Astrid Vanneste · 2025
Over the last years, intelligent systems have taken an increasingly large role in our lives. Machine learning (ML) is often applied in computer vision, speech recognition and for generating and processing language. There is also an enormous potential to use ML to automate the control of systems by using Reinforcement Learning (RL). For example, traffic light control can be improved by using an intelligent system to improve traffic flows through a city. For systems that consist of multiple controllable entities such as traffic lights, it is important to use methods that are designed to work with multiple entities. By using Multi-Agent RL (MARL), we can train intelligent agents that can improve the performance of multi-agent systems. In multi-agent settings, it can be useful to allow the agents to communicate with each other. For example, traffic lights can communicate to each other when large traffic streams are driving towards a certain traffic junction. However, it is often not straightforward to determine what information is important to share between agents. Therefore, communication between agents can also be learned. This way, the agents can determine by themselves which information is shared and will only communicate information that is useful for other agents. In this thesis, we look at MARL techniques and communication learning methods. More specifically, we investigate how to improve the efficiency and scalability of these techniques. We take a look at different parts of these systems that can be made more efficient or scalable. Most of our work is focused on improving the emergent communication between agents. In the first part of this thesis, we perform some exploratory research to investigate how emergent communication is affected when introducing a competitive aspect to the environment. Most existing communication learning methods, focus on fully cooperative environments. However, if we want to apply these methods in a wide range of applications, we need to look at more varied environments. As a first step, we look at emergent communication in a mixed cooperative-competitive setting. The second part focuses on improving the efficiency of the learned communication. First, we compare the mean operator and the attention mechanism as encoder mechanisms to summarize multiple messages into one encoding without losing important information. Next, we investigate a range of methods to represent the communicated information efficiently in the communication messages. Finally, we investigate the exploration strategy for MARL. State-of-the-art exploration methods choose to combine learning the behaviour to explore the environment (exploration) and the behaviour to reach the desired goal (exploitation). In this thesis, we investigate how we can train separate components for exploration and exploitation to be able to train agents that perform better in complex environments.