Notice of Removal: Anomaly Detection and Improvement of Clusters using Enhanced K-Means Algorithm

Vardhan Shorewala · 2021

This paper explores a unified approach to the improvement of clusters and the detection of anomalies in a dataset. It presents a novel approach for the formation of tighter and better clusters than traditional methods such as the K-means algorithm. It is evaluated against intrinsic measures for unsupervised learning such as the silhouette coefficient, Calinski Harabasz index, and the Davies Bouldin index. The proposed method decreases the intracluster variance of N clusters, until the variance approaches a global minimum. The method also extends K-means as a novel method for the detection of anomalies in the dataset. The proposed algorithm's performance is evaluated by extrinsic measures such as the Jaccard similarity score, the V-measure cluster, and the F1 Score. The algorithm is tested upon synthetic and real datasets, UCI Breast Cancer and UCI Wine Quality, to showcase its effectiveness in both cases. The proposed algorithm reduced the variance of the synthetic and real dataset, Wine Quality, by 18.7% and 88.1% respectively. It was also effective in increasing the accuracy and F1 score by 22.5% and 20.8% in the case of the Wine Quality dataset.

Read the paper · More papers on PaperTik