Research and Optimization of Distributed Clustering Algorithms for High-Dimensional Data
Runfeng Zou · Theory and Practice of Science and Technology · 2025
High-dimensional data analysis is a crucial task in various domains, including bioinformatics, image recognition, and text mining. However, traditional clustering algorithms, such as K-means and DBSCAN, struggle with the "curse of dimensionality," leading to inefficiencies in both accuracy and computational performance. To address these challenges, this paper proposes an optimized approach that integrates dimensionality reduction techniques with clustering algorithms in a distributed computing environment. Specifically, we evaluate and compare the effectiveness of principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), and autoencoders for dimensionality reduction, followed by K-means and DBSCAN clustering. Experiments are conducted on high-dimensional datasets, including text, image, and genomic data, to assess clustering quality, computational efficiency, and scalability. The results demonstrate that combining appropriate dimensionality reduction methods with clustering significantly improves performance while mitigating the impact of high dimensionality. Furthermore, our distributed computing framework enhances scalability, making it feasible for large-scale data processing. This study provides a comprehensive evaluation of different combinations of dimensionality reduction and clustering methods, offering practical insights into optimizing high-dimensional data clustering for real-world applications.