Hierarchical content classi cation and script determination for automatic document image processing

Zheru Chia, Wan-Chi Siua · 2003

Page segmentation and image content classi cation play an important role in automatic image processing with applications to mixed-type document image compression, form and check reading, and automatic mail sorting. In this paper, we rst present an enhanced background thinning based approach for fast page segmentation. After the analysis of three di4erent methods individually, a hierarchical approach for document content classi cation is proposed, which classi es a sub-image into one of two categories: text and halftone. Our approach combines a neural network model, cross-correlation metric, and Kolmogorov complexity measure in a hierarchical structure. Considering the necessity of a recognition system, we also propose using a three-layer feedforward neural network to classify text regions into Chinese and English scripts. The classi cation accuracy on a number of document images reaches 100% and 97.1% for halftone region and text region, respectively. Meanwhile, the system can achieve a correct rate of 92.3% and 95.0% for Chinese and alphabetic script determination, respectively. ? 2003 Pattern Recognition Society. Published by Elsevier Ltd. All rights reserved.

Read the paper · More papers on PaperTik