Interactive visualization system with multidimensional scaling

Min Song, Xia Lin · Proceedings of the American Society for Information Science and Technology · 2002

Visualization for analysis of textual data such as bibliographic data is now widely accepted by information analysts. As visualization has been a great interest of in the field of Information Science, a variety of techniques were adopted to visualize information. Multidimensional Scaling (MDS) is one of the major techniques and has led in recent years to a number of discoveries and insights. MDS is an iterative non-linear technique for projecting N-D data down to a lower number of dimension. N-D data points can, for example, be presented as 2-D display points. Although MDS was used heavily for visualization, to our knowledge, it was seldom to develop an interactive system available through Internet using MDS. It is partially attributed to the fact that MDS is slow to process large dataset and this slowness makes it impractical for MDS to be used in web environment. In this paper, we present an ongoing project associated with developing an interactive visualization system using MDS. To overcome the problem with MDS in processing data, we adapt a simulated annealing (SA) algorithm. SA is briefly described below. The system consists of three major components: 1) Conversion of input data, 2) Dimension Reduction, and 3) Display the results. The entire system was developed in Java. A prototype visualization system has been built around MDS. It allows a user to analyze data contained in appropriately formatted text files. Since it is cumbersome that the system only accepts the required format (Figure 1), in the next version of the system we plan to make the system flexible so that the user provide raw textual data and the system captures the meaningful set of vacabulries out of given data and make a proximity matrix. This is a new component of the system being developed, called Feature Extraction. Overview of the System The current format of raw textual data is as follows: Label and Term Frequency. Label is the first column of the matrix that represents the unit of the document. Term Frequency is represented in numerical value and starts from the second column to the n column. Each row represents a single entity in the dataset. So long as the format is preserved, any types of data such as e-mail, citation and web data are accepted. In Figure 1, the example of data is bibliographic data captured from Web of Science, which is a citation database provided by Thomson ISI. Lable is the first author's name and term freqency is calculated based on TF/IDF, which is widely used in IR. The input data is converted to the pairwise dissimilarities of the entities (in the example, a bibliographic record). Jaccard method is used to calculate similarities. Figure 2 is a visualization of a set of bibliographic data in the field of Knowledge Management. Based on the dissimilarity matrix, the system reduces dimension of the input data onto two dimensional space using a SA algorithm. Simulated annealing (SA) is an algorithm introduced by Metropolis et al. (1953) and is used to approximate the solution of very large combinatorial optimization problems. The technique originates from the theory of statistical mechanics and is based upon the analogy between the annealing of solids and solving optimization problems. A SA algorithm was used in a MDS-based visualization system, called MAVIS, developed by Bentley and Ward (1996). Visualization of the Field of Knowledge Management by the System After all the procedures are done, the multidimensional input data was mapped out onto two dimensional space. Visualization takes place in communication between Java Servlet and Applet. Data processing is done in the Servlet side and display of the results is based on Java applet. As illustrated in Figure 2, each node represents a bibliographic record. The label next to the node is the first author's name of the record. On clicking the node, the pop-up window containing full bibliographic information is brought up. As mentioned above, the goal of this paper is to report our ongoing project to develop an interactive system using MDS. This paper describes the architecture of the system briefly.

Read the paper · More papers on PaperTik