A system for intelligent document image analysis, recognition and compression

Wey-Wen Cindy Jiang · 1995

This dissertation presents a framework for solving tasks in intelligent document analysis systems. These systems aim at converting a document's pixel representation of text blocks into an equivalent symbolic representation and at removing redundancy in neighboring pixels of pictorial blocks. Document images are composed of blocks of different types, i.e. text, graphics, halftone images, and graytone images. We propose a document analysis system which performs segmentation, labeling, classification, and compression of blocks in document images. A flexible segmentation scheme employing three passes of run-length linking is designed. A block labeling scheme is used to label blocks of any shape. Our classifier can identify block types accurately by examining the texture of labeled blocks. Based on characteristics of each block type, an appropriate compression scheme is applied. Black and white graphics, which usually have long white runs or black runs, can be compressed with run-length coding. We propose a novel two-layer coding/decoding system to compress halftone blocks losslessly and efficiently. The multi-layer perceptron, a feed-forward neural network, is employed to exploit the non-linear correlation of neighboring pixels and to compress graytone images losslessly. Text images are recognized by optical character recognition (OCR) systems and stored in ASCII form. Most current OCR systems perform binarization on inputs before attempting recognition. However, a significant amount of information is lost during binarization of acquired graytone images. We propose a restoration scheme to convert graytone text images into bi-level images, while retaining information in graytones. Another good application of our restoration scheme is the recognition of texts in scene images. Our proposed framework provides several advantages over previous work: (1) Documents can have more complex layout styles. (2) Images segmented using our proposed scheme are better for classification. (3) Measures used for classification can more accurately discriminate among block types. (4) Our compression schemes for halftone blocks and graytone blocks outperform standard algorithms in compression ratio. (5) Text images, when restored into bi-level images using our proposed scheme, are closer to ideal ones and better adapted for character recognition. Applications of this work exist in any place where processing of printed information is needed.

Read the paper · More papers on PaperTik