Multi-Agent Reinforcement Learning for Moving Target Defense Decision-Making Based on Markov Game: An Adaptive Approach Under Three-Dimensional Metrics

Rongbo Sun, Jinlong Fei, Yuefei Zhu · 2025

Enhancing the effectiveness of cybersecurity defense requires not only advanced and practical defensive technologies but also relies on effective decision-making methods. In light of the complex and variable process of cyber offense and defense, the accurate, adaptive and effective selection of optimal strategies is a hot and challenging issue in the current research field of Moving Target Defense (MTD). We believe that existing MTD decision-making methods have significant deficiencies in accurately simulating the network environment, dynamically adapting to environmental changes, and balancing security, performance, and cost. Therefore, we propose a dynamic defense method that combines Markov game theory and improved Multi-Agent Reinforcement Learning (MARL). Firstly, we construct a Markov game model under complete information conditions to abstract the process of cyber offense and defense, considering the strategies of attackers and defenders and their corresponding state transitions to better fit the actual offensive and defensive game model. Secondly, we introduce three-dimensional indicators of security, performance, and affordability to construct the reward function of multi-agent reinforcement learning, thereby multidimensional quantifying the various effects of network defense strategies. On this basis, we design a reinforcement learning algorithm based on the multi-agent Q-learning network, which calculates the optimal network defense strategy measured by these three-dimensional indicators through the learning interaction process and introduces the concept of Probe-with-Penalty to ensure the stability of strategy updates and to achieve effective exploration and rapid convergence of the strategy space. Finally, we verify the effectiveness of the proposed model and algorithm through application examples and demonstrate the performance of the algorithm through result analysis. The study shows that the proposed model and algorithm can effectively improve the adaptability and efficiency of network defense, providing a new method for optimizing network defense strategies.

Read the paper · More papers on PaperTik