Knowledge acquisition and representation for document structure recognition: The CAROL Project

J. Schmidt, W. Putz · 2002

The authors describe a rule based recognition system to rebuild the structure of paper documents. This method is applied to an automatic cataloging system to be used in libraries. Documents are scanned and run through a character recognition engine. The result of the character recognition process is an output format with additional layout information serving as input for the rule interpreter. Rules for a specific document type are generated by a learning module, which enables the user to create a set of rules for a new document type. The learning component uses several generalization rules which can also be found in machine learning systems. CAROL, a demonstration prototype, is currently being tested by librarians.>

Read the paper · More papers on PaperTik