Scalable Data Stream Clustering with k Estimation

Paulo L. Candido, Murilo Coelho Naldi, Jonathan de Andrade Silva, Elaine R. Faria · 2017

The constant increasing of the generated data in real time has been creating new challenges for machine learning tasks and for one of their main branches: data clustering. The scenario of big data stream (high-speed data stream) has become reality. In order to deal with this scenario, new approaches are required. In this work, we present new techniques based on three clustering fields: Data Stream, MapReduce and automatic estimation of k from data. The goal is to cluster a high-speed data stream with varying number of clusters. Two scalable algorithms are proposed, based on centralized data stream algorithms. The first is based on the StreamKM++ and the second based on the F-EAC, an evolutionary algorithm used for batch clustering purpose. Results achieved the same high quality as the original centralized versions of the algorithms. The proposed techniques can be used instead the centralized ones when the velocity/volume is so high to fit in a centralized system.

Read the paper · More papers on PaperTik