Tabular Data Anomaly Detection Based on Density Peak Clustering Algorithm
Dong Liang, Jun Wang, Wenping Zhang, Yuqi Liu, Lei Wang, Xiaoyong Zhao · 2022 International Conference on Big Data, Information and Computer Network (BDICN) · 2022
Anomaly detection, also known as outlier detection, is one of the basic tasks of data mining, which aims to eliminate noise and discovers potential knowledge, and has received wide attention in the fields of data mining, machine learning, etc. Although many anomaly detection methods have been proposed, there are still some problems that have not been well addressed, including 1) Sample imbalance. Abnormal data only occupies a small portion of the entire dataset, making modeling in difficulty; 2) Lack of available labeled data; 3) Deep learning model has higher computational complexity. In this paper, we proposed TADPC(tabular anomaly density peak clustering) method, focus on tabular data, studying anomaly detection method based on density peak clustering algorithm, and introduce a method of pruning anomalies to ensure that the detected abnormal values are not the result of conventional noise, as a way to reduce false positives. Compared with the widely used anomaly detection algorithms such as ABOD(angle-based outlier detection), LOF(Local Outlier Factor), IForest(Isolation Forest), KNN(K- Nearest Neighbor) on six different datasets, our proposed method TADPC effectively improves the accuracy and reduces the false rate.