Research on Efficient Parallelization of Spectral Clustering Algorithm Based on Big Data
Chaoyue Wei · 2023
Cluster analysis is an important research direction in data mining. In the past decades, a large number of clustering algorithms have been developed, among which spectral clustering has been widely used because of its excellent performance in nonlinear separable data. In addition, with the continuous development of the Internet, more and more data are generated in the network, forming big data. Therefore, how to apply spectral clustering algorithm to big data and mine useful information has become a very important research topic. With the explosion of data scale, some industries have reached terabytes or even petabytes of data. The storage capacity and computing capacity of a single computer can no longer meet the requirements. At the same time, the spectral clustering algorithm will spend a lot of time in the process of processing massive data, which seriously limits the efficiency of the cluster analysis of mass data by the clustering algorithm in tomorrow. Therefore, improving the computational efficiency of the spectral clustering algorithm has become a research focus. In addition to using sampling method to optimize spectral clustering algorithm and reduce its computational complexity, with the emergence and wide application of various distributed parallel computing frameworks such as MapReduce and Spark, researchers have gradually focused on the distributed parallelization of spectral clustering algorithm. In the big data environment, the parallelization of spectral clustering algorithm is realized. It can not only obtain better clustering effect, but also improve the efficiency of spectral clustering algorithm for largescale cluster analysis. For the massive data of various industries, how to improve the large-scale clustering of spectral clustering algorithm. It is of great practical significance to analyze efficiency so as to improve work efficiency. According to the problem of long running time of spectral clustering algorithm in processing large-scale data sets, based on the idea of data parallelism, we study the efficient parallelization of spectral clustering algorithm based on Dask+CPUIGPU platform, which effectively improves the efficiency of spectral clustering algorithm in processing large-scale data sets.