Long range language models for free spelling recognition

Frank O. Thiele, Bernhard Rueber, Dietrich Klakow · 2002

Highly accurate spelling recognizers are essential in many commercially relevant applications. Examples are directory assistance, address taking in ordering services or help desks, and input of difficult or unknown words in dictation. Language models whose context length is flexibly configured by incorporating automatically determined letter groups code the structure of the recognition items in traditional bi- or trigrams. This gives a powerful method trading off computational demands versus recognition accuracy, comparing favorably with the standard word-list constraint. In addition, the automatic modeling of new words proves its benefits in applications with high out-of-vocabulary rates. For the task of spelling German last names over the telephone, letter error rates could be improved from 12.5% using a standard bigram to 3.6% with a trigram on a set of 2782 letter groups, giving the additional benefit of recognizing about 40% of the names not seen before in the language model training corpus.

Read the paper · More papers on PaperTik