Conditional random fields for multi-agent reinforcement learning
Xinhua Zhang, Douglas Aberdeen, S. V. N. Vishwanathan · 2007
Conditional random fields (CRFs) are graph-ical models for modeling the probability of labels given the observations. They have tra-ditionally been trained with using a set of observation and label pairs. Underlying all CRFs is the assumption that, conditioned on the training data, the labels are independent and identically distributed (iid). In this pa-per we explore the use of CRFs in a class of temporal learning algorithms, namely policy-gradient reinforcement learning (RL). Now the labels are no longer iid. They are ac-tions that update the environment and affect the next observation. From an RL point of view, CRFs provide a natural way to model joint actions in a decentralized Markov de-cision process. They define how agents can communicate with each other to choose the optimal joint action. Our experiments in-clude a synthetic network alignment problem, a distributed sensor network, and road traffic control; clearly outperforming RL methods which do not model the proper joint policy. 1.