Skew detection, page segmentation, and script classification of printed document images

B. Waked, Sabine Bergler, Ching Y. Suen, Sahar El Khoury · 2002

Automatic processing of international documents presents a number of challenging problems because Optical Character Recognition (OCR) techniques are not available for all languages and all script classes. Document images must be categorized according to their script type first, in our case Roman, Ideographic, or Arabic. We present a set of statistical methods that first detect and correct the skew of a document image. Next, the page is segmented into text and graphical components. The textual components are then segmented into paragraphs and lines; and finally we classify the script type into one of three categories. The system predicts the correct script category in 91% of cases when tested on real-life documents of varying kinds, diverse formats and qualities from many sources.

Read the paper · More papers on PaperTik