Parallelizing clustering of geoscientific data sets using data streams

Silvia Nittel, Kelvin T. Leung · 2004

Computing data mining algorithms such as clustering on massive geospatial data sets is still not feasible nor efficient today. In this paper, we introduce a k-means algorithm that is based on the data stream paradigm. The so-called partial/merge k-means algorithm is implemented as a set of data stream operators which are adaptable to available computing resources such as volatile memory and process-ing power. The partial data stream operator consumes as much data as can be fit into RAM, and performs a weighted k-means on the data subset. Subsequently, the weighted partial results are merged by a second data stream oper-ator. All operators can be cloned, and parallelized. In our analytical and experimental performance evaluation, we demonstrate that the partial/merge k-means can outper-form a one-step algorithm by a large margin with regard to overall computation time and clustering quality with in-creasing data density per grid cell. 1

Read the paper · More papers on PaperTik