Cluster forest based fuzzy logic for massive data clustering
Ines Lahmar, Abdelkarim Ben Ayed, Mohamed Ben Halima, Adel M. Alimi · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2017
This article is focused in developing an improved cluster ensemble method based cluster forests. Cluster forests (CF) is considered as a version of clustering inspired from Random Forests (RF) in the context of clustering for massive data. It aggregates intermediate Fuzzy C-Means (FCM) clustering results via spectral clustering since pseudo-clustering results are presented in the spectral space in order to classify these data sets in the multidimensional data space. One of the main advantages is the use of FCM, which allows building fuzzy membership to all partitions of the datasets due to the fuzzy logic whereas the classical algorithms as K-means permitted to build just hard partitions. In the first place, we ameliorate the CF clustering algorithm with the integration of fuzzy FCM and we compare it with other existing clustering methods. In the second place, we compare K-means and FCM clustering methods with the agglomerative hierarchical clustering (HAC) and other theory presented methods using data benchmarks from UCI repository.