Clustering Ensemble of Massive Data Based on Trusted Region

Suqin Ji, Ruowei Xing · 2021 3rd International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI) · 2021

Nowadays, there is a growing demand for clustering analysis of massive data, but the huge amount of data makes it a great challenge in storage and computing. Clustering ensemble forms a more robust result by ensemble multiple base clusterings, and the wrong labels in the base clusterings will greatly affect the ensemble result. This paper proposes a clustering ensemble method for massive data. Firstly, the massive data is divided into multiple small datasets by BLB sampling; K-means is used for clustering on each small dataset. In the base clusterings, find the trusted region according to a certain range from the cluster center. Combine all trusted regions to form the basic clustering structure of the data, and allocated all the remaining untrusted samples into the basic clustering structure to obtain the final clustering result. Experiments on synthetic datasets and real datasets show the effectiveness of the algorithm.

Read the paper · More papers on PaperTik