Classification by Ordering Data Samples (Acceleration and Visualization of Computation for Enumeration Problems)

Kazuya Haraguchi, Seok-Hee Hong, Hiroshi Nagamochi · Institutional Repositories DataBase (IRDB) · 2009

Visualization plays an important role as an effective analysis tool for huge and complex data sets in many application domains such as financial market, computer networks, biology and sociology.Howevcr, in many cases, data sets are processed by existing analvsis techniques (c.g., classification, clustcring, PCA) beforc applying visualization.In this paper, we study visual analysis of classification problem, a significant research issue in machine learning and data mining community.The problem asks to construct a $c1$ assifier from given sct of positive and negative samplcs that predicts the classes of future samples with high accuracy.We first extract a bipartite graph structure from the sample set, which consists of a set of samples and a set of subsets of attributes.We then propose ari algorithm that constructs a two-laycred drawing of tbe bipartite graph, by permuting the nodes using an edge crossing minimization technique.Thc resulting drawing can act as a new classifier.Surprisingly, experimental results on bench mark data scts show that our new classifier is competitive with a well-known decision trcc generator C4.5 in terms of prediction crror.Furthermore, the ordering of samples from tlic resulting drawing enables us to derive new analysis and insight into data such as clustering.high accuracy.Many cxisting methodologies arc based on geometric concepts and construct a hyperplanc as classificr.Classical ones (e.g.. Fisher's linear discriminant in $1930$ 's and perceptron in $1960' s[12]$ )

Read the paper · More papers on PaperTik