Robust models and novel similarity measures for high-dimensional data clustering

Duc Thang Nguyen · 2012

List of Figures x List of Tables xiiIn this thesis, we present our research works on some of the fundamental issues encountered in high-dimensional data clustering.We examine how statistics, machine learning and meta-heuristics techniques can be used to improve existing models or develop novel methods for the unsupervised learning of highdimensional data.Our objective is to achieve multiple key performance characteristics in the methods that we propose: reflecting the natural properties of high-dimensional data, robust to outliers and less sensitive to initialization, effective and efficient methods which are simple, fast, highly applicable and, on the other hand, produce good quality clustering results.Mixture Model-based Clustering, or M2C, is a clustering approach that has very strong foundation in probability and statistics.Among all the possible models, Gaussian mixture is the most widely used.However, when applied for very high-dimensional data such as text documents, it expresses a few disadvantages that do not exist in low-dimensional space.To explore and understand this matter thoroughly, an analysis of the impacts of high dimensionality to various aspects related to Gaussian M2C has been conducted.We propose an enhanced Expectation Maximization algorithm to help the Gaussian M2C go through the initialization stage more properly.Other than that, von Mises-Fisher is a kind of distribution coming from Directional Statistics and has recently been known as a suitable model for document data.Our application of the von Mises-Fisher distribution mixture as a Feature Reduction method shows interesting results in the document clustering problem.Experiments carried out on benchmarked document datasets confirm the performance improvements offered by the proposed methods.With the thesis, we also propose and present a novel clustering framework and the related algorithm to address the issue of clustering data with noise and outliers.The framework is called Partial Mixture Model-based Clustering, or PM2C.While the classical M2C framework does not take noisy data and outliers into consideration, the new framework is aware of the existence of these elements, vii PCM relaxes the constraint (2.14) to become u mi > 0, ∀m, i.It means membership degrees of an object to all the clusters must not sum to 1.However, PCM has its own drawback that it tends to produce overlapping clusters.An improved version of fuzzy-based clustering, called Possibilistic Fuzzy C-Means (PFCM), was introduced in [19].The authors combine two techniques into one in order to take advantage of each, and solve the problems of both.Three methods above serve as the basic background for fuzzy-based clustering approach.Nevertheless, they are still far from being efficient for document categorization.The intensive research work on this direction over the past decades has led to numerous variants of fuzzy-based clustering algorithm.Some of them are specifically designed for text clustering, such as Fuzzy Co-clustering of Documents and Keywords (Fuzzy CoDoK) in [20], Fuzzy Simultaneous KeyWord Identification and Clustering (FSKWIC) in [21], and Possibilistic Fuzzy Co-Clustering (PFCC) in [22].

Read the paper · More papers on PaperTik