ANALYSIS OF MACHINE LEARNING METHODS FOR AUTOMATING PENETRATION TESTING

Anastasiia Zhuravchak, Andrian Piskozub · Cybersecurity Education Science Technique · 2025

Automation of penetration testing using machine learning methods is one of the most promising areas in modern cybersecurity. The traditional approach to penetration testing requires significant resources, including financial ones, as well as the involvement of highly qualified specialists capable of conducting a comprehensive assessment of system security. This approach may not always provide sufficient speed in detecting new threats, especially in the face of the ever-increasing complexity of cyberattacks and the large number of vulnerabilities. The introduction of machine learning methods into the pentesting process allows creating flexible, adaptive systems that can not only automate routine tasks but also increase the accuracy and efficiency of vulnerability detection. This article provides an overview of the key machine learning algorithms that can be used to automate penetration testing, including support vector machines, random forest, naive Bayes, decision trees, and reinforcement learning methods. Each of these algorithms offers certain advantages in the context of vulnerability analysis, threat classification, and prioritisation of critical security issues. Special attention is paid to the role of large language models in the automation process. They can analyse logs, classify threats, generate reports, and even provide recommendations for fixing identified vulnerabilities. Such models can significantly increase the productivity of specialists by performing routine tasks automatically, which is especially useful when integrated with CI/CD processes. At the same time, the use of LLM has certain limitations, such as dependence on up-to-date data and high computing costs. The article also discusses the challenges and limitations of implementing machine learning algorithms in the pentesting process, such as the need for a large amount of high-quality data to train models, high computing resources, and the risks associated with possible false positives. The results of the study demonstrate that machine learning algorithms have significant potential to improve the efficiency of automated penetration testing, especially in large infrastructures with numerous vulnerabilities.

Read the paper · More papers on PaperTik