A Reinforced Active Learning Sampling for Cybersecurity NER Data Annotation

Smita Srivastava, Deepa Gupta, Biswajit Paul, Shubhashisa Sahoo · 2022

A vast majority of cybersecurity data comes in the form of unstructured textual data and needs to be annotated proficiently to train supervised machine learning models. The critical question is how much and which subset of data should be annotated for better model performance under budget constraints. Though most of the Machine Learning (ML) research focuses on learning better models using annotated datasets, this paper focuses on data annotation, specifically on suitable subset selection with an emphasis on Named Entity Recognition (NER) for cybersecurity. The proposed method provides an active learning based sampling strategy to select minimal yet most informative samples from a large set. Further, reinforcement learning is combined with the active learning approach to automate the process of sampling. The results on the auto-labelled cyber-NER dataset indicate that the cyber-NER model with Reinforced Active Learning (RAL) based sampling increases F1-Score by +2-7%and reduces compute time by 90% compared to random sampling based subset selection. Further, the proposed RAL approach achieved an 80% reduction in sample size and, consequently, annotation cost with comparable accuracy to that of complete selection.

Read the paper · More papers on PaperTik