5. Data Visualization
Society for Industrial and Applied Mathematics eBooks · 2007
Data visualization techniques are extremely important in cluster analysis. In the final step of data-mining applications, visualization is both vital and a possible cause of misunderstanding. This chapter introduces some visualization techniques. First, we introduce some nonlinear mappings, such as Sammon's mapping, multidimensional scaling (MDS), and self-organizing maps (SOMs). Then we introduce some methods for visualizing clustered data, including class-preserving maps, parallel coordinates, and tree maps. Finally, we introduce a technique for visualizing categorical data. 5.1 Sammon's Mapping Sammon Jr. (1969) introduced a method for nonlinear mapping of multidimensional data into a two- or three-dimensional space. This nonlinear mapping preserves approximately the inherent structure of the data and thus is widely used in pattern recognition. Let D = {x1, x2, …, xn} be a set of vectors in a d-space and {y1, y2, …, yn} be the corresponding set of vectors in a d* -space, where d* = 2 or 3. Let dis be the distance between xi and xs and dis* be the distance between yi and ys. Let y1(0), y2(0), …, yn(0) be a random initial configuration: y i (0) = ( y i1 (0) , y i2 (0) ,…, y id* (0) )T ,i=1,2,…,n. .