SELECTION OF THE NUMBER OF CLUSTERS IN K-MEAN ALGORITHM USING CLUSTER SOLUTION ENTROPY

V. I. Oreshkov · Vestnik of Ryazan State Radio Engineering University · 2021

The article discusses the problem of choosing the number of clusters in popular k-means clustering algorithm. It is noted that an unsuccessful choice of this hyper parameter can lead to the creation of a cluster structure the meaningful interpretation of which in the process of data mining leads to false conclusions and making incorrect management decisions based on them. The aim of the work is to develop a method for automatic selection of the number of clusters for k-means algorithm. The article provides an analytical review of the known methods for determining the number of clusters, their advantages and disadvantages being noted. The proposed approach is based on the elbow method, which uses the entropy of cluster solutions instead of the mean squares of clustering error. A practical example shows that the use of cluster solution entropy makes it possible to choose the number of clusters even in the case when the approach based on clustering error turns out to be untenable.

Read the paper · More papers on PaperTik