A novel threshold-based clustering method to solve K-means weaknesses
S. Ehsan Yasrebi Nayini, Somayeh Geravand, Ali Maroosi · 2017
Nowadays, using data mining techniques to discover hidden patterns in various data sets is very common. Among all data mining techniques, clustering algorithms are particularly important and K-means algorithm is one of the most popular clustering methods. Simplicity, flexibility and performance in large data sets are the most important advantages of the K-means algorithm. On the other hand, some factors such as determining the number of clusters by user, outlier data and sensitivity to initial centers of the clusters and thus possibility of reaching a local minimum can reduce the efficiency of the K-means algorithm. In the present study, the idea of the K-means clustering algorithm and its advantages and disadvantages is introduced and then a new method is proposed for data clustering which overcomes the limitations of the K-means algorithm. Therefore, accuracy and efficiency of the clustering method can strongly be increased. In the proposed algorithm, to avoid identifying the number of clusters by users, similarity and threshold measures are used for clustering. Also in this algorithm the outlier data are identified and hence their negative impact on the clustering can be avoided. The experimental results show that the proposed method can overcome K-means disadvantages and improve its efficiency.