Autoencoder Feature Residuals for Network Intrusion Detection: Unsupervised Pre-training for Improved Performance
Brian Lewandowski, Randy Clinton Paffenroth · 2022
Network intrusion detection is a constantly evolving field as researchers and practitioners work towards keeping up with novel attacks and growing amounts of network data. To aid in this challenge researchers have been exploring the use of deep learning techniques such as neural networks in order to detect zero-day attacks and reduce the amount of manual analysis required when a network intrusion detection system alert is generated. Herein we use an unsupervised pre-training step in order to take advantage of autoencoder feature residuals. We show that autoencoder feature residuals can be used in place of or in addition to an original feature set as input to a neural network classifier to improve classification performance. Often in such problems, experts perform feature engineering to optimize classification performance. However, such data manipulation is expensive and time consuming. Our novel approach provides a path that can alleviate the need for manual feature extraction while "doing no harm". That is, if the provided features are in some sense optimal, then our methodology will not degrade the classification performance. However, if the provided features are inefficient, then we demonstrate that our methodology can substantially improve classification performance on a broad range of benchmark cybersecurity datasets. Another practical side effect of using autoencoder feature residuals comes to light by analyzing the potential data compression benefits they provide.