READLEX: a lexicon for the recognition and analysis of structured documents
Rainer Hoch · 2002
This paper describes the architecture of a lexicon system called READLEX dealing with requirements of both text recognition and text analysis in document analysis. In order to meet these requirements, we have developed a concept for the automatic acquisition and generation of the lexicon. The heart of the lexicon system is based on redundant hash addressing techniques. Currently, the lexicon is used for the contextual post-processing of OCR results as well as the categorization of texts within structured documents. Other components for document analysis such as the address parser and a text pattern matcher also make use of the lexicon.