Thoughts on the Big Data Parallel Clustering Algorithm Based on the Mining of Community Maximal Classes
Hezheng Mao · 2025
In order to accurately and quickly find the network structure in big data, this paper proposes a big data clustering algorithm based on community maximal classes. To address the time consumption caused by the uncertainty of initial nodes and the calculation of the fitness function, local key nodes are introduced and the fitness formula is improved to reduce the time consumption. For the formation of the initial community, the concept of maximal clique is introduced. By analyzing the characteristics of maximal cliques, it is concluded that the core category of the community is composed of maximal cliques. Meanwhile, a method to obtain local core categories through the discovery of maximal cliques is proposed, and a parallel strategy for the maximal clique discovery algorithm is put forward. Then, the parallel strategy of the whole algorithm is proposed and experiments are conducted on real datasets. The experimental results prove that the algorithm proposed in this paper is feasible and effective, and is applicable to the discovery of network structures in large-scale data.