Test-Driven Development of Complex Information Extraction Systems using TextMarker.

Peter Klügl, Martin Atzmüller, Frank Puppe · 2008

Information extraction is concerned with the location of specific items in textual documents. Common process models for this task use ad-hoc testing methods against a gold standard. This paper presents an approach for the testdriven development of complex information extraction systems. We propose a process model for test-driven information extraction, and discuss its implementation using the rule-based scripting language TEXTMARKER in detail. TEXTMARKER and the test-driven approach are demonstrated by two real-world case studies in technical and medical domains.

Read the paper · More papers on PaperTik