Implementing Parallel Computing to Enhance the Performance of K-mean Algorithm

Ranyah Taha, Sara Alshakrani, Abdallah Alqaddoumi · 2021 International Conference on Data Analytics for Business and Industry (ICDABI) · 2021

Clustering is a common tool in data science, and it is used in many areas, such as statistics and bioinformatics. Clustering has a number of algorithms. The most well-known is k-means, which is popular due to its simplicity and efficiency. Similarly, as a programming model, the Message Passing interface (MPI4py) improves performance by reducing the execution time and increasing the speedup. By applying the K-means clustering algorithm to the MPI4py library, the performance of the algorithm in a parallel setting improved. The efficiency of clustering data using K-means is examined in this study in terms of overhead expenditure and execution between running the K-means algorithm sequentially and running the K-means algorithms using a Message Passing Interface (MPI) parallel architecture.

Read the paper · More papers on PaperTik