Network intrusion detection using word embeddings
Xiaoyan Zhuo, Jialing Zhang, Seung Woo Son · 2017
Word embeddings, learning syntactic and semantic relationships between words from the raw text, are known to achieve superior performance in many prediction and classification tasks. In this paper, we aim to show how word embedding techniques can work on a high number of nontextual data, thus has new applicability in cyber security. Specifically, we build neural embeddings with a large amount of network log data (KDD CUP'99 dataset) and train several classification models using the learned neural embeddings. Our experiment results demonstrate that the vector representations successfully learn useful features from non-textual data, similar to how classic word embeddings do with natural language and can achieve a high F1 score for both binary (normal vs. abnormal) and multi-class (DOS, Probe, U2R, R2L, and normal) classifications. We also demonstrate that the classification algorithms using the embedded vectors can maintain fairly high accuracy even when training data is very small.