Visualizing High-density Clusters in Multidimensional Data

Tran Van Long · 2010

The analysis of multidimensional multivariate data has been studied in various research areas for many years. The goal of the analysis is to gain insight into the specific properties of the data by scrutinizing the distribution of the records at large and finding clusters of records that exhibit correlations among the dimensions or variables. As large data sets become ubiquitous but the screen space for displaying is limited, the size of the data sets exceeds the number of pixels on the screen. Hence, we cannot display all data values simultaneously. Another problem occurs when the number of dimensions exceeds three dimensions. Displaying such data sets in two or three dimensions, which is the usual limitation of the displaying tools, becomes a challenge. The main approach consists of two major steps: clustering and visualizing. In the clustering step, we propose two clustering algorithms to construct hierarchical density clusters. In the visualizing step, we propose two methods to visually analyze the hierarchical density clusters. An optimized star coordinates approach is used to project high-dimensional data into the (two- or three-dimensional) visual space, in which the leaf clusters of hierarchical density clusters (well-separated in the original data space) are projected into visual space with minimizing the overlapping. The second method, we developed to visualize the hierarchical density cluster tree, combines several information visualization techniques in linked and embedded displays: radial layout for hierarchical structures, linked parallel coordinates, and embedded circular parallel coordinates. By combining cluster analysis with star coordinates or parallel coordinates, we extend these visualization techniques to cluster visualizations. We display clusters instead of data points. The advantage of this combination is scalability with both the size and dimensions of data set.

Read the paper · More papers on PaperTik