Three probabilistic language models for a large-vocabulary speech recognizer
Pierre Dumouchel, V. Gupta, Matthew Lennig, Paul G. Mermelstein · 2003
Relative performance is compared for three different language models applied to the linguistic decoding part of a 75000-word speech recognizer. These models are the trigram model, the tri-POS model (POS stands for parts of speech), and a smoothed trigram model with tied distributions for words three or more syllables long. The full trigram model gives the best performance but is most expensive in terms of data and storage requirements. The smoothed trigram and tri-POS models yield equivalent performance. For general text entry tasks, use of the tri-POS model is suggested since it is less sensitive to variations in the discourse domains. For applications specific to individual discourse domains, trigram models trained on data obtained from the target domain are recommended.>