Postprocessing algorithm based on the probabilistic and semantic method for Japanese OCR
Atsushi Konno, Yasuo Hongo · 2002
A postprocessing algorithm for Japanese OCR based on the probabilistic and semantic method is described. It determines the reliability of each recognized character and word in OCR outputs with their recognition probability and appearance frequency in the text. Using these reliabilities, grammatical word paths are searched in each phrase. When it is necessary to select the most suitable word from a similar word set, an attempt is made to select a particular one by the semantic method with co-occurrence word dictionary. This method was applied to OCR outputs of current newspapers and technical documents including some unregistered words, and evaluated its performance. While error correction rate depends on the ratio of unregistered words in the texts, the error detection rate is almost 90%.>