A Robust System Architecture for Mining Semi-structured Data

Lisa Singh, Bin Chen, Rebecca Haight, Peter Scheuermann, Kiyoko Flora Aoki-Kinoshita · 1998

The value of extracting knowledge from semi-structured data is readily apparent with the explosion of the WWW and the advent of digital libraries. This paper proposes a versatile system architecture for text mining that maintains structured data components in a relational database and unstructured concepts in a concept library. After a detailed explanation of our system architecture, we briefly describe IRIS, our prototype rule generation system Introduction Although much attention has been given to extracting knowledge from structured data, more and more tools that extract knowledge from semi-structured data are becoming available. The shift in focus is due in large part to the explosion of the World Wide Web (WWW) and the advent of digital libraries. Data from these arenas is potentially an invaluable source for analysis and decision support. The success of a text mining tool is dependent upon the ability to accurately represent document content and efficiently generate rules. This ...

Read the paper · More papers on PaperTik