Applying unsupervised grammar induction to OCR error correction
Samuel R. Scarano · 2014
This thesis presents a system for correcting errors from optical character recognition (OCR) software. As a noisy-channel error correction system, it uses a language model to provide a prior over the true text. For this purpose, we introduce a lexicalized version of Klein and Manning's Dependency Model with Valence, a grammar that is trained without structure annotation. The novel language model provides error correction performance that is slightly better than a 4-gram baseline on a corpus of historical English text. When interpolated with the 4-gram model, a relative reduction in word error rate of 32.1% is achieved, which is 2.5% more than the 4-gram model alone. However, the improvement is primarily attributable not to the modeling of dependency structure, but rather to the modeling of word classes, which is included therein. We determined this by achieving a similar improvement while constraining the model to a uniformly right-branching structure.