Keyword-Based Browsing and Analysis of Large Document Sets
Ido Dagan, Ronen Feldman, Haym Hirsh · 1996
Abstract Knowledge Discovery in Databases (KDD)focuses on the computerized exploration oflarge amounts of and on the discovery ofinteresting patterns within them. While mostwork on KDD has been concerned withstructured databases, there has been littlework on handling the huge amount ofinformation that is available only inunstructured textual form. This paperdescribes the KDT system for KnowledgeDiscovery in Texts. It is built on top of atext-categorization paradigm where textarticles are annotated with keywordsorganized in a hierarchical structure.Knowledge discovery is performed byanalyzing the co-occurrence frequencies ofkeywords from this hierarchy in the variousdocuments. We show how this term-frequency approach supports a range ofKDD operations, providing a generalframework for knowledge discovery andexploration in collections of unstructuredtext. Introduction Traditional databases store large collectionsof information in the form of structuredrecords, and provide methods for queryingthe database to obtain all records whosecontent satisfies the user's query. Morerecently, however, researchers in KnowledgeDiscovery in Databases (KDD) haveprovided a new family of tools for accessinginformation in databases (e.g. Brachman etal, 1993; Frawley et al, 1991; Kloesgen,1992; Kloesgen, 1995b; Ezawa and Norton,1995). The goal of KDD has been defined asthe nontrivial extraction of implicit,previously unknown, and potentially usefulinformation from given data (Piatetsky-Shapiro and Frawley 1991). Work in thisarea includes applying machine-learning andstatistical-analysis techniques towards theautomatic discovery of patterns in databases,as well as providing user-guidedenvironments for exploration of data.In Proceedings of the International Symposium on Document Analysis and Information Retrieval, 1996