Building Nonlinear Regression Models for Estimating the Number of Clusters and Their Initial Centroids

Sergiy B. Prykhodko, Natalia Prykhodko · 2023

The fundamental concept of partition-based clustering algorithms is to consider the data point’s center as the center of its associated cluster. In the initiation phase of clustering algorithms, notably the widely-researched k-means methodology, it is imperative to ascertain the number of clusters and establish their preliminary centroids (means). Now well-known methods define something one only, or the number of clusters, or initial centroids. We propose a technique for the simultaneous estimation of both the number of clusters and their initial centroids, which is based on building nonlinear regression models taking into account outliers. The use of nonlinear regressions is due primarily to the smaller width of their prediction intervals compared to the corresponding widths of linear regressions. Two application examples of the proposed technique using the four-variate Box-Cox normalizing transformation yield good results for the k-means clustering.

Read the paper · More papers on PaperTik