Automated penetration testing based on action clipping and decomposition

Qiankun Ren, Jingju Liu, Xinli Xiong · 2025

With the continuous development of network technologies, an increasing number of network devices are facing severe security threats, making it imperative to enhance the stability and reliability of network infrastructure. Automated penetration testing, which evaluates network security from the perspective of an attacker, faces challenges such as excessively large action spaces and insufficient utilization of historical information. To address these issues, this paper proposes a Hierarchical Reinforcement Learning Algorithm based on Action Decomposition and Clipping (HRLCD) to tackle the problem of poor decision-making performance in large-scale network environments. Firstly, HRLCD combines historical state representation with a curiosity-driven mechanism to optimize action selection strategies and improve the performance of the agent in penetration testing tasks. Secondly, the HRLCD algorithm employs an action clipping mechanism, which eliminates inefficient actions based on historical rewards, thus reducing the agent’s action space and enhancing decision-making efficiency. Finally, HRLCD introduces a hierarchical strategy framework that divides the decision-making process into highlevel and low-level strategies, with the former responsible for global decisions and the latter for specific action execution. This hierarchical structure not only improves task processing efficiency but also enhances the system’s adaptability and flexibility. In the POMDP-based decision model for penetration testing, this method effectively accelerates the training process and improves the accuracy of the policy. Experimental results show that the HRLCD algorithm significantly outperforms traditional methods in decision-making performance in penetration testing tasks, providing new insights for the application of reinforcement learning in the field of network security.

Read the paper · More papers on PaperTik