Decentralized Reinforcement Learning with Risk Aversion in Multi-Agent Systems
Daichi ISHIKAWA, Taisei Ichino, Naoki Hayashi, Masahiro Inuiguchi · 2025
This study addresses risk-averse distributed reinforcement learning in multi-agent systems, where agents collaboratively optimize their policies with a risk index in the objective function. Unlike conventional reinforcement learning, which maximizes the expected cumulative reward, the proposed risk-averse reinforcement learning incorporates Conditional Value at Risk (CVaR) to account for rare but significant adverse events. We propose a risk-averse distributed reinforcement learning algorithm based on the policy gradient method, which enables agents to collaboratively update their policies by sharing information through a communication network. Through a theoretical analysis, we show that the proposed algorithm ensures agreement among agents on their estimated policies. Moreover, we show that the consensus value converges in a neighborhood of a locally optimal solution, which ensures that each agent’s learned policy remains aligned with risk-aware optimization criteria.