An Approach to Text Mining using Information Extraction
Haralampos Karanikas, Christos Tjortjis, Babis Theodoulidis · 2000
In this paper we describe our approach to Text Mining by introducing TextMiner. We perform term and event extraction on each document to find features that are likely to have meaning in the domain, and then apply mining on the extracted features labelling each document. The system consists of two major components, the Text Analysis component and the Data Mining component. The Text Analysis component converts semi structured data such as documents into structured data stored in a database. The second component applies data mining techniques on the output of the first component. We apply our approach in the financial domain (financial documents collection) and our main targets are: a) To manage all the available information, for example classify documents in appropriate categories and b) To "mine" the data in order to "discover" useful knowledge. This work is designed to primarily support two languages, i.e. English and Greek. 1.