Document Image Analysis For Reading Books

Yoshitake Tsuji, Jun Tsukumo, Ko Asai · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1987

A fundamental problem in machine vision is to detect and identify special objects in an image. In the field of machine-reading for existing printed matter and books, a very important technique allows extracting and recognizing characters in desired text lines from a document image. This paper describes a hierarchical image segmentation, which separates a document image into its entities. Furthermore, a character segmentation, with minimum variance criterion, and a character recognition, based on three improved loci feature, have been developed as two elemental methods for reading books. In these experimental results using different commercial Japanese pocket books, 99% of text lines were correctly extracted. Also, it was successful in reading 99.30% of the Japanese characters and Chinese ideographs, as used in printed text.

Read the paper · More papers on PaperTik