Japanese OCR Error Correction Using Stochastic Morphological Analyzer and Probabilistic Word N-gram Model
Koichi Takeuchi, Yūji Matsumoto · International Journal of Computer Processing Of Languages · 2000
While the accuracy of current OCR systems is getting very high, they are still error-prone. In this paper, we clarify how much of recognition errors in text can be corrected using linguistic information from on-line texts. We present an OCR error correction method which uses character trigram, stochastic morphological analysis and word trigram models. These models are trained on a large untagged text. The proposed method does not use any graphical information about characters. Therefore the method can be applied to any domain that has a large on-line text corpus. When our method is applied to text which include random character substitution, it improves a text of 90% correct character rate into that of 94.3% correct rate and a 95% correct text into a 96.9% correct one.