d-FuzzStream: A Dispersion-Based Fuzzy Data Stream Clustering
Leonardo Schick, Priscilla de Abreu Lopes, Heloisa A. Camargo · 2018
Fuzzy clustering algorithms have recently been investigated as appropriate techniques to extract knowledge from Data Streams due to their unsupervised nature and flexibility to deal with changes in the distribution of data. While most fuzzy clustering algorithms for Data Streams are based on chunks, the FuzzStream algorithm, proposed before by the authors of this paper, pioneered a fuzzy extension of a different approach known as the Online-Offline Framework (OOF). The extended framework, named Fuzzy Online-Offline Framework (FOOF), includes two steps known as fuzzy abstraction and fuzzy clustering. The fuzzy abstraction step continuously summarizes data in a set of cluster features called Fuzzy Micro Cluster (FMiC). Then, these FMiCs are later clustered in the fuzzy clustering step to generate the data partition. Although FuzzStream has shown to be more robust than other OOF-based algorithms, the fuzzy abstraction process in the algorithm overly reduces the data summarization, almost producing one FMiC for each example, also suffering from high overlapping FMiCs. Furthermore, the algorithm has a long processing time due to its need to calculate membership matrices for every example. In this paper we propose the d-FuzzStream algorithm, an adaptation of FuzzStream using the concepts of fuzzy dispersion and fuzzy similarity in order to improve the data summarization while minimizing the complexity of the algorithm. Experiments showed that the proposed algorithm generates FMiCs with higher representativeness and lower execution time than its original version, still producing similar clustering results.