Robust language-independent OCR system
Zhidong Lu, Issam Bazzi, András Kornai, John I. Makhoul, Prem Natarajan, Richard M. Schwartz · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1999
We present a language-independent optical character recognition system that is capable, in principle, of recognizing printed text from most of the world's languages. For each new language or script the system requires sample training data along with ground truth at the text-line level; there is no need to specify the location of either the lines or the words and characters. The system uses hidden Markov modeling technology to model each character. In addition to language independence, the technology enhances performance for degraded data, such as fax, by using unsupervised adaptation techniques. Thus far, we have demonstrated the language-independence of this approach for Arabic, English, and Chinese. Recognition results are presented in this paper, including results on faxed data.