Research on Intrusion Detection Based on Feature Extraction of Autoencoder and the Improved K-Means Algorithm

Xingang Wang, Linlin Wang · 2017

Nowadays, intrusion detection is a technology to effectively avoid a number of risks of network intrusion. The K-means algorithm is widely used in intrusion detection. But, the algorithm has some shortcomings, such as random selection of k value, sensitive selection of initial cluster centers, and low accuracy in clustering high-dimensional data. In order to make up for these shortcomings of the K-means algorithm, this paper proposes an AE-Kmeans architecture which combines an autoencoder with the improved K-means algorithm. The AE-Kmeans architecture realizes the dimension reduction and feature extraction of these original data by introducing an autoencoder, and uses the improved K-means algorithm to cluster these processed data. These improvements of the improved K-means algorithm mainly include two aspects: the first aspect, a new method is introduced to select initial cluster centers; the second aspect, the algorithm calculates the weight of each attribute with the coefficient of variation, and then uses these weights in Euclidean distance formula, resulting in a weighted Euclidean distance formula. Finally, the AE-Kmeans architecture uses KDD CUP99 data set for intrusion detection simulation experiments. The Experimental results show that the AE-Kmeans architecture not only enhances the ability to deal with high-dimensional data, but also improves the detection rate and reduces the error-detection rate compared with K-means algorithm.

Read the paper · More papers on PaperTik