Risk-averse multi-agent reinforcement learning with distributional mean-variance formulation

Jiahui Jiang, He Wang, Wenwu Yu · 2025

This paper explores the development of risk-sensitive policies for cooperative multi-agent reinforcement learning within dynamical environments. The notion of risk refers to the uncertainty of the accumulated reward. By measuring the value of risk, agents can acquire distributions of accumulated rewards, facilitating a more precise estimation of Q values. In this paper, we propose a novel MARL method employing distributional mean-variance formulation for addressing risk-averse multi-agent tasks. The value function of each agent is approximated under the mean-variance risk measure with distribution of returns for decentralized execution. During the early stages of training process, a soft-risk level scheduling is designed for efficient exploration with the measure of risk. Ultimately, the risk-related policies are optimized by utilizing mean-variance values as target estimators in the TD error through centralized training. Empirical results demonstrate that the superiority of our method over state-of-the-art MARL methods in the designed multi-intersection TSC problems, showcasing improved coordination and stability.

Read the paper · More papers on PaperTik