Initializing the EM Algorithm for Data Clustering and Sub-population Detection

Zhengyu Hu · OhioLink ETD Center (Ohio Library and Information Network) · 2015

In this thesis our research on two loosely related parts are presented: initialization for the EM algorithm for model-based data clustering, and sub-population detection.Clustering is the task of putting objects into groups in such a way that the observations in the same group are more "similar" to each other than those in other groups.Clustering is a useful tool for summarizing the data, and hence could greatly reduce the complexity of a data set.Moreover, data could be compressed using clustering so that it take less space to store.There is no universally accepted definition of cluster, so there exist many different clustering algorithms.Among these algorithms, model-based clustering using the Gaussian mixture model is built on sound mathematical foundation, and is widely applied in many areas.However, the prevailing method for finding the maximum likelihood estimator (MLE), i.e., the expectation maximization (EM) algorithm, is very sensitive to initialization.Hence, the EM algorithm is very likely to stuck in a local maximum of the likelihood function, especially when the number of clusters in the data is large, but the problem of initializing the EM algorithm is not very well studied.My deep and sincere gratitude first and foremost goes to my former advisor, Dr.

Read the paper · More papers on PaperTik