Evaluating various Finetuned LLMs for Cybersecurity Named Entity Recognition
Shanmukha Aditya G, B. S. Kruthika, Venkata Sneha G, Deepa Gupta, Smita Srivastava · Procedia Computer Science · 2025
In an era dominated by escalating cybersecurity threats, the rapid and accurate identification of named entities within textual data is crucial. Named Entity Recognition (NER) serves as a cornerstone in this endeavor, facilitating the extraction of vital information from unstructured sources. Our study explores the development of a robust NER model tailored for cybersecurity text analysis, leveraging advanced deep learning techniques and state-of-the-art Large Language Models (LLMs). Utilizing an English Cyber Security NER dataset with 24 tags, we interrogate the performance of diverse fine-tuned LLMs, including Mistral 7B, BART 406M, FLAN T5 Small 580M, and FLAN T5 Base 2.1B. By evaluating key metrics such as accuracy, precision, recall, and F1-score, we assess each model’s efficacy in domain-specific entity recognition. Our research demonstrates an achieved F1 score of approximately 74% after 100 training epochs, highlighting the potential of these models in enhancing cybersecurity defenses. The comparative analysis of these fine-tuned LLMs provides insights into their respective strengths in tackling cybersecurity NER tasks, contributing to the ongoing efforts in identifying and mitigating threats effectively.