DTC: Addressing the long-tailed problem in intrusion detection through the divide-then-conquer paradigm

Chaoqun Guo, Nan Wang, Yuanlin Sun, Dalin Zhang · 2023

Intrusion detection systems (IDS) analyze the monitored data to detect patterns or signatures that correspond to known cyberattack techniques, vulnerabilities, or deviations from established baselines. They employ different algorithms and techniques to identify potential threats. When building "Known Patterns and Signatures", it always faces the long-tailed problem, which refers to the imbalanced distribution of different types of network traffic or events in a dataset, which means that there are very few instances of certain types of intrusions compared to more common types of network traffic or benign events. Such a situation poses a great challenge to deep learning-based or machine learning-based detection models on how to handle it. Models may struggle to learn from long-tailed distributions because they tend to bias their predictions toward the majority, and perform well on common events but poorly on rare intrusions. To address this problem, different from previous methods concerning training the model on the whole samples to obtain a balanced data distribution, we first focus on dividing the whole samples into a balanced group and an imbalanced one, then, we train a detection model on the balanced group. Cycle over and over again. Specifically, We propose to use a Gaussian mixture flow filter to progressively perform sample aggregation, continuously transforming the long-tail distribution into a more balanced which allows us to train the classifier on the obtained balanced group. The separated training samples with high distribution balance make it easier to train subsequent classifiers and mitigate the head-to-tail bias. Through extensive experiments, we have achieved new state-of-the-art performance on common intrusion detection datasets such as UNSW-NB15, CIC-IDS2017, and NSL-KDD. These results demonstrate that it is possible to surpass carefully constructed balanced datasets by progressively distinguishing the head class and the tail class.

Read the paper · More papers on PaperTik