Efficient Data Stream Clustering Algorithm Based on k-Means Partitioning and Density
Weiwei Ni, Lu Jieping, Geng Chen, Sun Zhi-hui · Mini-micro Systems · 2007
Data stream clustering is an important issue in data stream mining. Most of the existing algorithms adopted K medians (means) method to solve this problem, which are not suitable to address the problem of clustering high dimensional or abnormal distributed data streams. This article proposes a k-Means partitioning and density based data stream clustering algorithm—CLUSMD. The algorithm applies K means clustering on each partition of the data stream to generate mean reference point set, and subsequently density based clustering is applied to these reference points to get the clustering result of each periods. Theoretic analysis and experimental results showe that CLUSMD is effective and efficient.