Classification and distribution of optical character recognition errors

Jeffrey Esakov, Daniel Lopresti, Jonathan S. Sandberg · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1994

This paper describes an approach for classifying OCR errors based on a new variation of a well-known dynamic programming algorithm. We present results from a large-scale experiment we performed involving the printing, scanning, and OCRing of over one million characters in each of three fonts. Times, Helvetica, and Courier. Our data allows us to draw a number of interesting conclusions about the nature of OCR errors for a particular font, as well as the relationship between error sets for different fonts.

Read the paper · More papers on PaperTik