Fuzzy and possibilistic co-clustering techniques for high-dimensional data analysis
William Chandra Tjhi · 2008
In this thesis, we present our study of some new data clustering mechanisms for highdimensional data analysis.Our objective is to simultaneously achieve several goals for data analysis namely: effective clustering in high dimension, rich and natural representations of clusters, robustness to outliers, and highly-interpretable clusters.Novel clustering models based on the fuzzy and possibilistic co-clustering frameworks have been developed for the purpose.Co-clustering, a technique of simultaneous clustering of objects and features, is regarded by many in the literature to be one of the most effective approaches to automatically categorize high-dimensional data.Fuzzy co-clustering enhances standard crisp co-clustering by capturing a more realistic fuzzy representation of co-clusters, i.e. the resulting associated object and feature clusters.Existing prominent fuzzy co-clustering algorithms however are vulnerable to outliers.In addition, the common mathematical model that underlies these existing algorithms can be shown to be too rigid in formulation.This rigidity prevents possible expansions of fuzzy co-clustering that can potentially address the outlier problem, as well as enrich the representations of co-clusters.We introduce a new fundamental and less rigid formulation for fuzzy co-clustering based on the dual-partitioning approach.In essence, the dual-partitioning approach is like the Fuzzy C-means in the context of fuzzy co-clustering, and it offers a solid ii