Weak model-dependent page segmentation and skew correction for processing document images

J.F. Cullen, Koichi Ejiri · 2002

Presents an algorithm for fast accurate page segmentation that is as far as possible independent of a model for the document page. The method is based on a reduced image data representation that uses bounding rectangles and run length size distributions contained within the bounding rectangles. These rectangles are the basis for the method of skew detection, column identification and merging of text into blocks, and they help achieve accurate page segmentation. The reduced complexity that rectangles offer insures a fast processing time. The method is applied to a broad range of documents found in a typical office environment. Documents are scanned at 400 dpi and stored as binary images.>

Read the paper · More papers on PaperTik