Clustering in multivariate data: visualization, case and variable reduction

Sunhee Kwon · 1999

Cluster analysis is a very common problem for multivariate data. It is receiving intense attention due to the current boom in data warehousing and mining driven by the growth in information technology today. Technology is allowing us to collect massive data sets, both in cases and variables, and develop sophisticated interactive and dynamic graphics. There are three current issues for cluster analysis: visualizing cluster structure, reducing the number of cases, and reducing the number of variables in very large data sets. This thesis addresses each of these issues;The lower-dimensional projection of data found by projection pursuit which preserves the cluster structure helps clustering by eliminating the influence of nuisance variables. Initially partitioning data into a set of small classifications improves the efficiency of hierarchical agglomerative clustering by saving the time and memory for the beginning stage of clustering. Minimal spanning tree is used for this partitioning method.

Read the paper · More papers on PaperTik