K-means Clustering for Large Data: Anomaly Detection in Supervisory Control and Data Acquisition Systems

Corwin Stanford, April L. Tanner · 2023

K-means is a common unsupervised method for partitioning low dimensional datasets. However, it is usually not used for higher dimensional datasets, such as those found in Supervisory Control and Data Acquisition (SCADA) systems. Here we examine its application to a large dataset for the purposes of detecting cyber attacks using anomaly-based intrusion detection in the Battle of the Attack Algorithms (BATADAL) dataset. Additionally, dimensionality reduction using Principal Component Analysis (PCA) is examined as a method of improving anomaly detection performance. Emphasis is placed on methods for selecting parameters without using information obtained from examination of attack labels.

Read the paper · More papers on PaperTik