UAVThreatBench: A UAV Cybersecurity Risk Assessment Dataset and Empirical Benchmarking of LLMs for Threat Identification

Padma Iyenghar · Drones · 2025

UAVThreatBench introduces the first structured benchmark for evaluating large language models in cybersecurity threat identification for unmanned aerial vehicles operating within industrial indoor settings, aligned with the European Radio Equipment Directive. The benchmark consists of 924 expert-curated industrial scenarios, each annotated with five cybersecurity threats, yielding a total of 4620 threats mapped to directive articles on network and device integrity, personal data and privacy protection, and prevention of fraud and economic harm. Seven state-of-the-art models from the OpenAI GPT family and the LLaMA family were systematically assessed on a representative subset of 100 scenarios from the UAVThreatBench dataset. The evaluation applied a fuzzy matching threshold of 70 to compare model-generated threats against expert-defined ground truth. The strongest model identified nearly nine out of ten threats correctly, with close to half of the scenarios achieving perfect alignment, while other models achieved lower but still substantial alignment. Semantic error analysis revealed systematic weaknesses, particularly in identifying availability-related threats, backend-layer vulnerabilities, and clause-level regulatory mappings. UAVThreatBench therefore establishes a reproducible foundation for regulatory-compliant cybersecurity threat identification in safety-critical unmanned aerial vehicle environments. The complete benchmark dataset and evaluation results are openly released under the MIT license through a dedicated online repository.

Read the paper · More papers on PaperTik