Internal-Cluster-Validation-Based Model Selection For k-Means Clustering

Mozammel H. A. Khan · 2024

Clustering is an important data mining process that partitions the data points of an unlabeled dataset based on the similarities of the data points. The k-Means algorithm is a widely used clustering method. We propose a model selection method for k-Means clustering algorithm that identifies the number of optimal clusters and selects the suitable feature subset for better clustering of a dataset. For this purpose, we use internal cluster validation index called Sum of Euclidean Distances (SED). To make the SED Index produced by different feature subsets comparable, we preprocess the dataset and define a per feature SED Index. We generate the candidate feature subsets using a Genetic Algorithm and choose the feature subset that produces the minimum per feature SED Index. We use the per feature SED Index for determining the optimal number of clusters. We also propose a method for selecting initial centroids of the clusters for better clustering performance. We validate our proposed model selection method using Iris, Wine, and Seeds datasets from UCI Machine Learning Repository.

Read the paper · More papers on PaperTik