Density Micro-Clustering Algorithms on Data Streams: A Review

Amineh Amini, Teh Ying Wah · 2011

Abstract—Data streams are massive, fast-changing, and infinite. Applications of data streams can vary from critical scientific and astronomical applications to important business and financial ones. They need algorithms to make a single pass with limited time and memory. Mining data streams is concerned with extracting knowledge structures represented in models and patterns in non-stopping data streams. Clustering is a prominent task in mining data streams, which group similar objects in a cluster. Several clustering algorithms have been introduced in recent years for data streams that are based on distance, so they can find only spherical shapes. Therefore, density-based clustering algorithms are adopted for data streams with ability for not only discovering the arbitrary shape clusters, but also for providing protection against the outliers. In fact, in density-based clustering algorithms, dense areas of objects in the data space are considered as clusters, which are segregated by low density area (noise). However, in the clustering data streams, due to certain characteristics, it is impossible to record all the data. Micro-clusters are a technique in stream clustering that maintains the compact information about the data objects in data streams. Microcluster is a temporal extension of the cluster feature, which compresses the data effectively. In this paper, we intend to review the outstanding density-based clustering algorithms on data streams using micro-clusters. We will explore algorithm characteristics and analyze their merits and limitations.

Read the paper · More papers on PaperTik