Advanced state clustering for very large vocabulary HMM-based on-line handwriting recognition

Andreas Kosmala, Daniel Willett, Gerhard Rigoll · 1999

The paper presents some novel methods for the introduction of context dependent hidden Markov models (HMM) to online handwriting recognition. The use of these so-called n-graphs can lead to substantially improved modeling accuracy, but requires some intelligent parameter reduction methods (state clustering). This is especially the case for the investigated very large vocabulary system, incorporating an active vocabulary of 200000 words. Switching from context independent models to context dependent models-considering the underlying vocabulary-yields in the worst case to 25000 HMMs and very poor trainability for most of the introduced models. Therefore, the conducted investigations are focused on an appropriate state clustering method which is supported by decision trees and some new self organizing approaches to generate the required trees. The presented comparison takes also the different context dependencies (left, right or both sides) into consideration.

Read the paper · More papers on PaperTik