Trigram-based algorithms for OCR result correction
Konstantin B. Bulatov, Temudzhin Manzhikov, Oleg Slavin, Igor Faradjev, Igor M. Janiszewski · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2017
In this paper we consider a task of improving optical character recognition (OCR) results of document fields on low-quality and average-quality images using N-gram models. Cyrillic fields of Russian Federation internal passport are analyzed as an example. Two approaches are presented: the first one is based on hypothesis of dependence of a symbol from two adjacent symbols and the second is based on calculation of marginal distributions and Bayesian networks computation. A comparison of the algorithms and experimental results within a real document OCR system are presented, it's showed that the document field OCR accuracy can be improved by more than 6% for low-quality images.