N-Gram Language Models for Offline Handwritten Text Recognition

Matthias Zimmermann, Horst Bunke · 2004

This paper investigates the impact of bigram and trigram language models on the performance of a hidden Markov model (HMM) based offline recognition system for handwritten sentences. The language models are trained on the LOB corpus which is supplemented by various additional sources of text, including sentences from additional corpora and random sentences produced by a stochastic context-free grammar (SCFG). Experimental results are provided in terms of test set perplexity and performance of the corresponding recognition systems. For the text recognition experiments handwritten material from the IAM database has been used.

Read the paper · More papers on PaperTik