LLM-CVX: A Benchmarking Framework for Assessing the Offensive Potential of LLMs in Exploiting CVEs

Mohamed Amine El yagouby, Abdelkader Lahmadi, Mehdi Zakroum, Olivier Festor, Mounir Ghogho · 2025

The increasing capabilities of Large Language Models (LLMs) in code generation and reasoning have raised concerns about their potential misuse, particularly for automating or assisting with vulnerability exploitation tasks. This concern highlights the need for a systematic evaluation of the offensive potential of LLMs. Existing methodologies in this context use synthetic vulnerabilities, rely on fixed prompting strategies and exploit tools in their evaluation, and do not consider efficiency. In this paper, we introduce LLM-CVX, a novel benchmarking framework designed to systematically evaluate LLMs on real-world CVE exploitation tasks, with extensibility towards prompting strategies and exploit tools, and allowing LLMs to make multiple attempts to capture both effectiveness and efficiency. We implemented this framework to evaluate 14 state-of-the-art LLMs on exploiting 36 real CVEs using 2 exploit tools (Metasploit and GitHub PoC) and a correction loop that enables LLMs to correct their previous exploitation attempts. Experimental results reveal variation in the behavior of different LLMs, using both exploitation tools, with closed-source models generally outperforming open-source ones.

Read the paper · More papers on PaperTik