Evaluating Three Large Language Models for Security Incident Detection and Analysis

Abdullah Mudzakir, Jip Tyrone Ativiirya Janaprasetya, Emanuel Kayowuan, Ventje Jeremias Lewi Engel, Bagaskoro Saputro · Procedia Computer Science · 2025

As modern cyberattacks become more complex, traditional security tools are often not enough, so there is a need for smarter tools to help with Digital Forensics and Incident Response (DFIR). This study addresses that need by comparing three 8-billion-parameter Large Language Models, such as Dolphin3 8B, LLaMA3.1 8B, and Qwen3 8B, to see how effective they are in helping with DFIR tasks. Using a setup that included real Apache2 attack logs and a special knowledge base of 133 documents from sources like NIST and OWASP, the models were tested through a Retrieval-Augmented Generation (RAG) system. Their performance on forensic tasks, such as identifying attack stages and extracting Indicators of Compromise (IOCs), was measured against a ground truth from GPT-4o using Cosine Similarity and BERTScore. The results clearly show that Qwen3 8B, when helped by the knowledge base, performed much better than the other models, receiving the highest average score of 0.7233. It also achieved the top scores for both Cosine Similarity (0.6128) and BERTScore (0.8338). In time, this research shows that while LLMs have great potential in cybersecurity, not all models benefit from extra knowledge, so there isn’t a single solution that works for every case. This highlights that choosing the right model is very important; a model like Qwen3 8B, when used correctly with a RAG system, can greatly improve the accuracy of forensic analysis, proving to be a valuable tool for Blue Team operations.

Read the paper · More papers on PaperTik