An intelligent text extraction and navigation system

Jakub Piskorski, Günter Neumann · 2000

We present sppc, a high-performance system for intelligent text extraction and navigation from German free text documents. sppc consists of a set of domainindependent shallow core components which are realized by means of cascaded weighted finite state machines and generic dynamic tries. All extracted information is represented uniformly in one data structure (called the text chart) in a highly compact and linked form in order to support indexing and navigation through the set of solutions. German text processing includes (among others) compound processing, high performance named entity recognition and chunk parsing based on a divide-and-conquer strategy. sppc has a good performance (4380 words per second on standard PC environments) and high linguistic coverage. 1 Introduction The information society will gradually provide its members with nearly unlimited access to all sorts of information. Without highly e#ective sophisticated information extraction applications for deali...

Read the paper · More papers on PaperTik