Highly degraded text recognition in the framework of hidden markov models
H. E. Meadows, Chinching Yen · 1996
In this dissertation, an optical character recognition system with new training and recognition approaches is presented. The goal is to achieve high recognition performance over highly-degraded and connected document images. Based on the Pseudo two-dimensional Hidden Markov Models, this system is built to directly recognize gray-scale document images. Information loss caused by the binarization process is thus avoided. Some issues such as automatic generation of initial models for new applications and compensation for gray-level images of different scanning quality have also been tackled. Traditionally, the Viterbi decoding process has been a popular approach to find the optimal match between observation and models. Various path duration information can be incorporated during the decoding process to improve the results. Due to the lack of a complete path map, the advantage of this incorporation cannot be fully accessed. We therefore propose the duration-corrected N-best hypotheses search to improve the decoding process. It is a backward tree search initiated upon the completion of the forward Viterbi process. During this backward search, both the complete forward path map from the Viterbi pass and the partial path information incrementally collected during the backward pass are combined to impose more accurate duration constraints. The duration-corrected optimal match at the end of the backward search gives improvement over the traditional Viterbi result. Multiple hypotheses are also found through this procedure for postprocessing to get even higher recognition rates. A new training scheme is then introduced with the use of these N-best hypotheses. It improves models trained by traditional Maximum Likelihood (ML) criteria. The objective of the popular ML method is to reach a set of model parameters such that the likelihood function over the training set could be maximized. Our new scheme, however, builds a direct link between error rate and model parameters. The objective of the parameter optimization is now to directly minimize the recognition error rate instead of maximizing the likelihood function value. It also exposes models with competitive word hypotheses that do not exist in the original training set. This new scheme thus improves recognition rates and model robustness over the traditional ML approach.