Big Data Clustering Algorithms: Improving Efficiency and Scalability for Massive Data Mining and Pattern Recognition
Shanmugam Muthu, G. Ramesh Kalyan, V. Samuthira Pandi, D Shobana, Lakshmi Priya J, S Jaishika · 2025
The emergence of big data has completely changed the way data analysis is done. As a result, there is a pressing need for effective clustering algorithms that can handle large datasets from different fields. Efficiency, scalability, and finding meaningful patterns in enormous data are some of the important difficulties that this research tackles in relation to big data clustering. We thoroughly examine and evaluate a range of clustering algorithms, from more conventional ones like hierarchical and K-Means clustering to more cutting-edge ones like density-based spatial clustering (DBSCAN) and model-based clustering. To further improve the performance and scalability of clustering algorithms, we investigate new methods that make use of distributed computing frameworks like Apache Spark and MapReduce. We offer a thorough assessment of their efficacy in terms of clustering quality, computational efficiency, and resource use by using and evaluating these methods on large-scale datasets. While conventional clustering algorithms may fail miserably when faced with extremely big data sets, our research shows that newer methods that make use of optimization techniques and distributed processing greatly outperform their predecessors. The effect of feature selection and dimensionality reduction techniques on improving clustering results is also shown. By shedding light on how to choose and implement clustering algorithms for effective data mining and pattern identification, this study adds to the continuing endeavors in the domain of big data analytics. Progress in areas like customer segmentation, anomaly detection, and social network analysis can be achieved with the help of adaptive algorithms that can handle large datasets while keeping clustering results accurate and dependable, according to the study.