Use of synthesized images to evaluate the performance of optical character recognition devices and algorithms
Frank R. Jenkins, Junichi Kanai · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1994
Synthesizing document images is a cost effective way to create a large test database and allows researchers to control typesetting and noise variables. Yet the effectiveness of using synthesized images in optical character recognition (OCR) research has not been extensively investigated. In this project, three kinds of test databases were used to study the performance of OCR devices: digitized `real world' documents, page images synthesized from ASCII files, and the synthesized images printed and digitized. The cleanest synthesized images were not necessarily recognized most accurately. Our results suggest that, in addition to typographical features and noise, linguistic features affect the performance of an OCR device.