Penetration Test Path Discovery Based on NHSC-PPO
Fang Li, Xueyan Wang, Yining Liu, Li Luo · 2023
Aiming at the problems that Deep Q Network (DQN) and its variant Deep Reinforcement Learning (DRL) algorithm are not efficient and difficult to converge in the penetration test attack path discovery scenario, this study introduces the Normalized Entropic Dynamic Reward Scaled Clipped-Proximal Policy Optimization Algorithm (NHSC-PPO) for determining the optimal attack path. This algorithm enhances the Proximal Policy Optimization (PPO) algorithm, and introduces advantage function normalization, policy entropy, dynamic reward scaling, and gradient clipping techniques. We first convert the network topology and vulnerability information into an attack graph by MulVAL, and then train the Agent to continuously explore and update the strategy on the attack graph by using the NHSC-PPO algorithm, and finally find the attack path with the highest reward value in the dataset. Experimental results in various scenarios show that, compared with PPO, Double Dueling DQN and DQN algorithms, the penetration testing path discovery of NHSC-PPO algorithm can more efficiently and accurately converge to the path with the highest reward, and then determining the optimal attack path.