A Framework for Semi-Automatic Development of Rule-based Information Extraction Applications

Peter Kluegl, Martin Atzmueller, Tobias S. Hermann, Frank Puppe · LWA · 2009

For the successful processing and handling of (large scale) document collections, effective information extraction methods are essential. This paper presents a framework for the semiautomatic development of rule-based information extraction applications based on the TEXTMARKER language utilizing machine learning methods. We describe the approach in detail and present the TEXTRULER system as an implementation of the proposed approach.

Read the paper · More papers on PaperTik