Analyzing the Parallel Computing Performance of Unsupervised Machine Learning
Vishnu Vardhan Baligodugula, Fathi Amsaad, Noor Zaman Jhanjhi · 2024
The popularity of unsupervised machine learning techniques is increasing drastically as they can generate clusters of data samples. It facilitates making critical decisions and overcoming challenges in various applications. In the realm of data analysis, clustering algorithms are commonly used to cat egorize large datasets into smaller groups. As datasets continue to grow in size, traditional clustering methods become impractical, and parallel computing has become an attractive option to improve performance. The proposed method investigates the use of unsupervised learning to classify data and determine the GDP of countries as low, medium, or high. Three clustering techniques are compared for their performance, scalability, and ease of use on large datasets: Minibatch K-Means, parallel Minibatch K-Means using Massage Passing Interface, and Parallel Minibatch K-Means using the cloud. The results highlight the importance of parallel computing in data analysis and demonstrate the effectiveness and efficiency of parallel clustering techniques.