CPOCEDS-Concept Preserving Online Clustering for Evolving Data Streams
K. T. Jafseer, S. Shailesh, A. Sreekumar · Research Square · 2022
Abstract Clustering streaming data is challenging due to many temporal dynamics,such as concept drift, concept evolution, and feature evolution.Concept evolution is the most challenging of these. Due to concept evolution,new classes may emerge or existing classes may disappear, soit is crucial to process streaming data continuously. This paper proposesa novel online clustering method, specifically for streaming datawith concept evolution. It consists of three phases: initialization, clusteringand outlier handling. To identify recurrences of previous datain streaming data, it is critical to preserve the sequential propertiesof data chunks. In the proposed model, representatives from previouswindows are added to the current window, making it distinct from existingmodels. The detection and handling of outliers are very challengingtasks in streaming data analysis. Outliers are often the first instancesof a new cluster. The proposed model stores the outliers from eachdata window. When the number of outliers exceeds a certain threshold,the representatives of outliers are added to the next window toidentify new classes. To handle the lack of data sets for the training ofsuch models, we created a synthetic data set with 22020 data instances.Using Silhouette Coefficient, Calinski-Harabasz Index, and Davies-Bouldin Index analysis, this model yielded the most favourable results.