User-Defined Expected Error Rate in OCR Postprocessing by Means of Automatic Threshold Estimation

J. Ramón Navarro-Cerdán, Joaquim Arlandis, Juan-Carlos Pérez-Cortés, Rafael Llobet · 2010

In this work, a method for the automatic estimation of a threshold that allows the user of an OCR system to define an expected error rate is presented. When the OCR output is post-processed using a language model, a probability, a reliability index (or a “transformation cost”) is usually obtained, reflecting the likelihood (or its inverse) that the string of OCR hypotheses belongs to the model. Using a threshold on this index (or cost) to reject the less reliable hypotheses, a variable level of expected accuracy can be imposed on the output. It is much more convenient for the user the ability to “fix” at an acceptable level the expected error rate instead of having to deal with an arbitrary threshold. Of course, the result will always be high reject rates for difficult tasks and lower reject rates for easier tasks.

Read the paper · More papers on PaperTik