A Clustering Technique for Mining Data from Text Tables

Hasan Davulcu, Saikat Mukherjee, I. V. Ramakrishnan · 2002

Considerable quantities of valuable data about product information and financial statements is often available in sources and formats that are not amenable for querying using traditional database techniques.One such important source is text documents.In such documents these kinds of data often appear in tabular form.A data item in these text tables may span several words (e.g.product description).Furthermore items supposedly within the same column do not necessarily begin or end at the same position.Thus the absence of any regularity in column separators makes it difficult to automatically mine, i.e. extract data items from text tables.Nevertheless an interesting characteristic often exhibited by these tables is that intra-column items are "closer" to each other than inter-column items.We exploit this observation to develop a clustering-based technique to extract data items from these tables.In contrast to previous appproaches, a unique and important aspect of using clustering is that it makes the technique robust in the presence of misalignments.We provide a characterization theorem for text tables on which this technique will always produce a correct extraction.We discuss the design and implementation of a system for extracting tabular data based on this clustering technique.We present experimental evidence of its effectiveness and usability on real industrial data.

Read the paper · More papers on PaperTik