Particle-based language modelling

E. W. D. Whittaker, Philip C. Woodland · 2000

This paper investigates the use of particle (sub-word) N-grams for language modelling. One linguistics-based and two datadriven algorithms are presented and evaluated in terms of perplexity for Russian and English. Interpolating word trigram and particle 6-gram models gives up to a 7.5% perplexity reduction over the baseline word trigram model for Russian. Lattice rescoring experiments are also performed on 1997 DARPA Hub4 evaluation lattices where the interpolated model gives a 0.4% absolute reduction in word error rate over the baseline word trigram model. 1. INTRODUCTION Most of the current approaches to language modelling for speech recognition tend to use words, or classes of words, as the modelling units. Words are a logical choice, since it is ultimately words that are to be output by a speech recognition system, but they are not necessarily the best units for capturing dependencies in a text. The optimal set of units will inevitably depend on the language, the sparsity of the ...

Read the paper · More papers on PaperTik