Design of a Multi-Agent Collaborative Decision-Making System Based on Reinforcement Learning
Zhe Wei, Meiying Zhou · International Journal of Pattern Recognition and Artificial Intelligence · 2025
This paper proposes a modular multi-agent reinforcement learning framework that integrates Centralized Training with Decentralized Execution (CTDE), attention-based communication, and adaptive reward shaping. Built upon an extended Soft Actor–Critic algorithm, the system enables decentralized agents to learn robust policies under partial observability. A shared critic computes value estimates using the full global state, while decentralized actors use selectively aggregated peer messages via a task-driven attention mechanism. Adaptive reward shaping dynamically aligns agent incentives with global objectives, accelerating convergence. The system is evaluated on three benchmarks: Multi-Agent Particle Environment (MPE), StarCraft II Micromanagement Challenge (SMAC), and a custom Resource Allocation Simulator (RAS). Compared to MAPPO, MADDPG, and ISAC baselines, our method improves average episodic reward by 15–25%, reduces convergence steps by up to 40%, and enhances coordination scores significantly. Results also show superior stability across random seeds and reduced wall-clock training time, highlighting the method’s effectiveness for real-world deployment in dynamic multi-agent settings.