Multimodal LLMs for Zero-Shot Intrusion Detection Using NetFlow Visualisations
Majed Luay, Siamak Layeghy, Yash Pandey, Gayan K. Kulatilleke, Marius Portmann · 2025
This paper presents a proof-of-concept framework integrating time-windowed NetFlow traffic visualisations with zero-shot inference from multimodal LLMs for network intrusion detection. Real-world network traffic, augmented with manually injected attacks such as IP sweep, DoS, and DDoS, is transformed into scatter plots representing host communication structure and traffic volume for each 1-minute window. These visualisations are assessed by GPT-4o and LLaVA in a zero-shot setting, without task-specific fine-tuning or prior examples. Results show that multimodal LLMs, particularly GPT-4o, effectively detect structural anomalies in network traffic using visual patterns alone, especially when prompts are enhanced with detailed explanations of attack appearances in the visualisations. These findings underscore the potential of LLM-based visual intrusion detection for NetFlow data and encourage further research into domain-specific visual encodings, optimised prompt design, and multimodal LLM adaptations for network intrusion detection applications.