On Some Data Pre-processing Techniques For K-Means Clustering Algorithm
Dauda Usman, Fatima Sani Stores · Journal of Physics Conference Series · 2020
Abstract This paper analyzed the performance of the basic K-Means clustering algorithm with two major data pre-processing techniques and superlative similarity measure with automatic initialization of seed values on the dataset. Further experiment was conducted with simulated data sets to prove the accuracy of the new method. The new method presented in this paper gave a good and promising performance for the different types of data sets. The sum of the squares clustering errors reduced significantly for the new method as compared with basic K-Means method whereas inter-distances between clusters are preserved to be as large as possible for better clusters identification.