Data clustering for very large datasets plus applications

Tian Yao Zhang · Minds at UW (University of Wisconsin) · 1997

Data clustering is an important way of exploring data, and has been shown to be useful in many domains such as data classification and image processing. Recently, there is a growing emphasis on exploratory analysis of very large datasets to discover useful patterns. It is called data mining, and data clustering is regarded as a particular branch. However existing data clustering methods do not adequately address the problem of processing very large datasets with limited resources (e.g., running time and memory). As the dataset size increases, they do not scale up well in terms of memory requirement, running time, and result quality. In this thesis, with a new in-memory data structure called CF-tree serving as a data distribution summarization, an efficient and scalable data clustering method is proposed, and implemented in system BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies). With various patterns of synthetic datasets, its performance is studied and compared with other methods in terms of memory requirement, running time, quality, stability and scalability. Finally, BIRCH is applied to solve two real world problems: (1) building an iterative and interactive pixel classification tool, and (2) generating initial codebook for image compression. Its performance on these real datasets is also compared with other methods.

Read the paper · More papers on PaperTik