An expectation-maximization algorithm working on data summary

Huidong Jin, Kwong‐Sak Leung, Man–Leung Wong · 2002

Scalable cluster analysis addresses the problem of processing large data sets with limited resources, e.g., memory and computation time. A data summarization or sampling procedure is an essential step of most scalable algorithms. It forms a compact representation of the data. Based on it, traditional clustering algorithms can process large data sets efficiently. However, there is little work on how to effectively make cluster analysis on data summaries. From the principle of the general expectation-maximization algorithm, we propose a model-based clustering algorithm to make better use of these data summaries in this paper. The proposed EMACF (Expectation-Maximization Algorithm on Clustering Features) algorithm employs such data summary features as weight, mean, and variance explicitly. It is proved that EMACF converges to a local maximum likelihood value. EMACF is linear with the number of data summaries instead of data items, and thus can be integrated with any efficient data summarization procedure to construct a scalable clustering algorithm. I.

Read the paper · More papers on PaperTik