Exploring the Potential of Large Language Models for Red Teaming in Military Coalition Networks
Erik Adler, Johannes F. Loevenich, Linnet Moxon, Tobias Hürten, Florian Spelter, Johannes Braun, Yann Gourlet, Thomas Lefeuvre, Roberto Rigolin F. Lopes · 2024
This paper reports on an ongoing investigation comparing the performance of large language models (LLMs) in generating penetration test scripts for realistic red agents. The goal is to develop human-level adversaries (red agents) in an automated cyber operations gym environment instrumented to train a team of blue agents. Our methodology defines five approaches for structuring the prompts used to generate Metasploit scripts to exploit vulnerabilities described in Common Vulnerabilities Exposures (CVEs). These approaches have been tested using three LLMs, namely GPT-4o, WhiteRabbitNeo, and Mistral-7b. GPT-4o is used as the baseline for our comparison study. The results suggest that GPT-4o outperforms the other LLMs in all experiments. However, the results also suggest that Mistral-7b can be fine-tuned to achieve acceptable performance while consuming much less computational and memory resources during execution due to a smaller number of parameters: 7 billion parameters for Mistral-7b versus 1.76 trillion for GPT-4o.