Evaluating Large Language Models for Anomaly Detection with Cost-Efficient Sampling: A Generalized Framework
Duy Anh Pham, Ievgeniia Kuzminykh, Andrii Astrakhantsev, Bogdan Ghita · 2025
Cybersecurity threats, particularly zero-day attacks, pose significant risks to organisations. Traditional machine learning (ML) models struggle to adapt to novel threats outside of their training data. Large Language Models (LLMs), with their adaptability and semantic reasoning, offer potential for detecting anomalies in network traffic. This paper assesses whether LLMs can effectively support anomaly detection when deployed in industrial contexts with real-time constraints and large-scale data streams. To lower computational costs, we propose two-level stratified sampling and optimised prompt engineering to reduce data volume by over 80%, while still preserving key malicious patterns. Comparing LLaMA in zero-shot and few-shot modes to ML baselines using the CICIDS2017 dataset, we find LLMs fall short on volumetric attacks. We discuss strategies to improve LLMs for cost-effective production use.