A Learning-Based POMDP Approach for Adaptive Cyber Defense Against Multi-Stage Attacks
Yuantian Zhang, Weixia Cai, Huashan Chen, Zhenyu Qi, Hong Chen, Feng Liu, Sen He · 2024
While various defense mechanisms have been proposed in cybersecurity, it is still unclear how these defense mechanisms should be dynamically employed to mitigate the damage of multi-stage attacks. In this work, we consider the problem of generating defense strategy in real-time to thwart multi-stage attacks. We use the Bayesian condition dependency graph (BCDG) to model the interactions between the attacker and the defender. Considering that both the attacker and the defender have uncertainty about their respective observations, we formulate the strategy selection problem as a partially observable Markov decision process (POMDP), where the attacker and the defender need to find their optimal strategies in a partially observation environment. To solve the problem of state space explosion, we develop a deep reinforcement learning (DRL) based approach to seek the optimal strategies. We conduct experiments with various settings to evaluate the effectiveness of our approach. Experiment results show that our DRL-based approach outperforms baselines, and the approach is robust to the uncertain security environment.