High-dimensional Data Clustering and Statistical Analysis of Clustering-based Data Summarization Products
Dunke Zhou · OhioLink ETD Center (Ohio Library and Information Network) · 2012
With the advancement of modern technology, we have seen the expansion of data in two dimensions: number of variables and number of observations.Such highdimensionality and large data volume have posed new challenges to statistical analysis.This thesis considers two problems related to cluster analysis: high-dimensional data clustering and statistical analysis of clustering-based data summarization products.High-dimensionality often makes traditional clustering methods ineffective.Variable selection is a common approach to reduce the dimensionality of data for better cluster analysis.Most of recently developed methods either explicitly or implicitly perform variable selection based on variable importance (V I) measure.In this thesis, an algorithmic framework is introduced which iterates between constructing V I and performing variable selection conditioning on each other.Within this framework, we develop an ensemble V I which is constructed by averaging a set of V I s.Both theoretical and simulation studies show that the proposed ensemble V I has better variable selection performance than unensemble V I and is robust to the choice of the number of groups in cluster analysis.In addition to the development in V I, we propose a new V I-based variable selection method which selects a set of variables through sequentially testing the existence of group structure in data.Its effectiveness is demonstrated through simulation study and a real data application.My deep and sincere gratitude first and foremost goes to my advisor, Dr.