DATA SETS FOR OCR AND DOCUMENT IMAGE UNDERSTANDING RESEARCH

Isabelle Guyon, Robert M. Haralick, Jonathan J. Hull, Ihsin T. Phillips · WORLD SCIENTIFIC eBooks · 1997

Several signi cant sets of labeled samples of image data are surveyed that can be used in the development of algorithms for o ine and online handwriting recognition as well as for machine printed text recognition. The method used to gather each data set, the numbers of samples they contain, and the associated truth data are discussed. In the domain of o ine handwriting, the CEDAR, NIST, and CENPARMI data sets are presented. These contain primarily isolated digits and alphabetic characters. The UNIPEN data set of online handwriting was collected from a number of independent sources and it contains individual characters as well as handwritten phrases. The University ofWashington document image databases are also discussed. They contain a large number of English and Japanese document images that were selected from a range of publications.

Read the paper · More papers on PaperTik