Navigating High-Dimensional Data with Advanced Clustering Algorithms
Ziyan Shirin Raha, M. Khandaker, Md. Zahid Hossain, Shamim Akhter · 2025
This research presents a comprehensive analysis of various high-dimensional clustering algorithms applied to a large dataset comprising 5,000,000 data points. To manage the high dimensionality, Principal Component Analysis (PCA) reduces the datasets high dimensionality while keeping substantial variance. The study evaluates various clustering algorithms: DBSCAN, HDBSCAN, Spectral Clustering, BIRCH, Autoencoder + KMeans, KMedoids and CLARANS. Performance is assessed using Silhouette Score, Davies-Bouldin Score, and Calinski-Harabasz Index, measuring cluster compactness, separation and dispersion ratio, respectively. Our findings demonstrate that BIRCH excels with a high Silhouette Score of 0.9685, the lowest Davies-Bouldin Score of 0.0650, and a Calinski-Harabasz Index of 2105.4370, proving its ability in constructing well-defined, cohesive clusters. This study combines quantitative metrics and visual analyses to assist discover the best clustering approaches for high-dimensional data.