A Comparative Study and Analysis of Dimensionality Reduction Techniques on High Dimensional Datasets for Network Anomaly Detection

Mettu Jhansi Rani, Dhanpratap Singh · 2024

Finding network anomalies is an important part of keeping computer networks safe and secure. As network data gets more complicated and bigger, its multidimensionality has become a big problem for programs that look for different patterns. By cutting down the number of features in high-dimensional datasets, dimensionality reduction methods play a major role in solving this problem. This study discusses about several methods that are available for reducing the number of dimensions, such as Principal Component Analysis (PCA), Autoencoders, Linear Discriminant Analysis (LDA), and t-Distributed Stochastic Neighbour Embedding (t-SNE). We test these methods on several large network datasets, such as the KDD Cup 1999 dataset and the NSL-KDD dataset. One of the factors for assessment is how well the methods can lower the number of dimensions in the datasets while keeping the important traits related to network errors. Based on the features of the dataset, PCA performed well at lowering the number of dimensions in a dataset while keeping most of the variance. It might not work well for finding non-linear connections in the data, though. On the other hand, t-SNE and Autoencoders are better at recording the relationships that are not linear but they may be hard to compute.

Read the paper · More papers on PaperTik