Contrastive Trajectory Learning for Multi-Agent Reinforcement Learning Policy Transfer

Yu Wang, Quan Liu, Hao Chen, Ke Fu, Linyue Liu, Benke Gao, Xiyao Ding, Jian Huang · 2025

Cooperative multi-agent reinforcement learning (MARL) has achieved significant success in various applications. While parameter sharing is commonly used in existing methods to improve training efficiency and learn task-specific policies, it often leads to homogeneous behaviors among agents. This homogeneity severely limits the policy's exploration capability within the task's state space. Consequently, when these policies are transferred to new tasks, this poor exploration frequently results in suboptimal learning performance. To address this challenge, we propose a Contrastive Trajectory Learning (CTL) method to enhance policy generalization across different tasks. Specifically, we use an attention-based entity disentanglement mechanism to accurately identify key entities within an agent's observations, preventing irrelevant entity information from being embedded into the trajectory representation. In addition, we incorporate a permutation-equivariant network to adapt to the dynamic mapping relationship between observation entities and the action space across tasks. This decouples the policy distribution and improves the sample efficiency of policy transfer. Moreover, to boost policy exploration over different task space, we introduce a contrastive disagreement (CD) loss between the trajectory representations of different agents to learn discriminative and diverse trajectory representations. Experimental results on the RealSim Empowered Learning Arena (RELA) demonstrate that CTL achieves efficient transfer across different tasks.

Read the paper · More papers on PaperTik