Automatic generation of custom document image decoders

Gary E. Kopec, Philip A. Chou · 2002

A framework for document image recognition, called document image decoding (DID), that supports the automatic generation of custom document recognition systems from user-specified document models is discussed. A document recognition problem is viewed as consisting of a message source, an imager, a noisy channel, and an image decoder (recognizer). The inputs to a decoder generator are explicit models for the message source, imager and channel; the output is a specialized program that decodes an image in terms of these models. The models used in DID are based on a stochastic attribute grammar model of document production. Use of an automatically generated decoder to analyze telephone yellow pages is described.>

Read the paper · More papers on PaperTik