Auto-K K-Means: Automated Cluster Selection for High-Dimensional Data Using the Second Derivative Elbow Method

S Udhayasuriyan, S. Saradha · 2025

The selection of the optimal number of clusters ($K$) is a critical challenge in unsupervised learning, particularly in high-dimensional feature spaces. Traditional K-Means clustering requires manual selection of$K$, leading to subjectivity and potential inaccuracies. This paper proposes an Auto-K K-Means clustering framework that integrates deep feature extraction with an optimized Elbow Method based on second derivative curvature detection. The methodology consists of dataset preprocessing, feature extraction using a pre-trained deep learning model, automatic determination of the optimal number of clusters, and clustering evaluation using quantitative metrics. The proposed method is evaluated on benchmark datasets, achieving an improvement in clustering compactness, with a Davies-Bouldin Index (DBI) of 10.80, comparable to Fixed K-Means (DBI = 9.54) while eliminating manual$K$selection. The consistently low Silhouette Scores across all methods highlight challenges in highdimensional clustering, emphasizing the need for dimensionality reduction. The findings demonstrate the effectiveness of the AutoK K-Means method in reducing manual intervention while maintaining competitive clustering performance, making it suitable for high-dimensional data applications.

Read the paper · More papers on PaperTik