LB-MDP-CL: A Reinforcement Learning and Co-Evolutionary Approach for Optimizing Responses in Multi-Step Cyberattacks

Aws Naser Jaber, Giordano Colò · 2025

Cyberattacks pose significant threats to business-critical operations, necessitating robust and adaptive defense strategies. This paper introduces a Learning-Based Markov Decision Process with Coevolutionary Learning (LB-MDP-CL) framework to optimize defensive responses in multi-step cyberattack scenarios. By simulating repeated interactions between defenders and evolving attacker models, the proposed approach enables continuous learning and adaptation, thereby enhancing operational resilience and overall effectiveness. We present detailed simulations, calibration methodologies, and a reproducible setup for evaluating LB-MDP-CL. Furthermore, we compare LB-MDP-CL with established Reinforcement Learning (RL) methods, demonstrating the potential benefits of co-evolutionary learning in dynamic cyber defense. Experimental results show that LB-MDP-CL significantly reduces attacker dwell time while balancing risk reduction and system continuity.

Read the paper · More papers on PaperTik