Visual and analytical tools for record level cluster analysis
Urška Cvek, Georges Grinstein · 2004
Data generation and acquisition in a variety of scientific fields has been exponentially increasing over the past decade. Exploratory analysis techniques try to find a path through this vast space of data, creating groups or hierarchies, finding outliers, patterns and meaningful domain-specific insights. One knowledge discovery technique, cluster analysis, is concerned with the “natural” grouping of records into clusters without establishing the rules to separate the records into categories (the role of classification). A plethora of clustering algorithms exists and each has its own interpretation of the definition of a cluster or group of records. No single clustering method or approach exists today that would be the uniformly recommended method, not even for a particular data domain. This dissertation presents exploratory and interactive visual and analytical tools and techniques that aid the comparisons and analyses of clustering results from multiple algorithms and parameter settings. Our cluster analysis has been elevated from the level of a cluster to the level of a record. We present new record-based analytical record similarity measures and their impact on record, cluster and data cluster analysis and separation. Two visual methods for exploration of the new measures are the stable matrix and the cluster tree approach. A formal platform model has enabled the development of the Cluster Comparator tool set that includes the analytical and visual tools, all developed within the framework. The tools are enriched and guided by domain specificity and we focus on questions and problems in the biomedical and functional genomic domains.