Modeling Markov Decision Process Elements Leveraging Language Models for Offline Multi-Agent Reinforcement Learning

Zhuohui Zhang, Bin He, Bin Cheng · 2024

Offline multi-agent reinforcement learning (MARL) has attracted considerable attention for its capability to learn distributed policies from offline datasets, eliminating the need for costly exploratory interactions or accurate environmental models. However, the absence of real-time interactions can result in inaccuracies when learning Markov Decision Process (MDP) elements, thereby affecting sample efficiency. Drawing inspiration from the relational learning capabilities of language models, we propose leveraging these models to implicitly capture MDP elements, such as global state representations. This paper introduces the state language model (SLM) framework, a sequence-to-sequence approach that implicitly learns global state representations from offline data. The SLM framework enables agents to infer global-like information by modeling implicit relationships between local observations and the global state, without relying on predefined communication protocols. A feedback mechanism further refines state predictions, ensuring both scalability and accuracy. Extensive experiments on offline MARL benchmarks demonstrate that SLM can be seamlessly integrated into existing algorithms, significantly enhancing their performance.

Read the paper · More papers on PaperTik