Long-Term Performance of a Generic Intrusion Detection Method Using Doc2vec

Mamoru Mimura, Hidema Tanaka · 2017

Cyber attack techniques in cyberspace are evolving every second, and detecting unknown malicious communication is a challenging task. Pattern-matching-based techniques and using malicious website blacklists are not effective, because the attacker can easily change the traffic pattern or the attack infrastructure. There are many behavior based detection methods which use the characteristic of drive-by-download attacks or C&C traffic. However, many previous methods specialize the attack techniques and the adaptability is limited. Moreover, they have to decide the feature vectors every attack method. Accordingly, we propose a generic detection method which is independent of attack methods and does not need devising feature vectors. This method uses Paragraph Vector an unsupervised algorithm that learns fixed-length feature representations from variable-length pieces of texts, such as sentences, paragraphs, and documents, and learns the context in proxy server logs. This paper reveals the long-term performance and the effectiveness to the other dataset. This paper conducts timeline analysis with the dataset which contains captured traffic from Exploit Kit (EK) between 2014 and 2017. This paper also demonstrates cross-dataset validation by showing that an automated feature extraction scheme learned from one dataset can be used successfully for classification on another dataset. The experimental results show the proposed method is effective over three years, and effective on another dataset too. The proposed method achieves an F-measure of 0.95 in the timeline analysis and an F-measure of 0.96 on the other dataset.

Read the paper · More papers on PaperTik