Evaluating pruned k-TSS Language Models: perplexity and word recognition rates
Amparo Varona · 2010
A syntactic approach based on regular grammars, the k-Testable in the Strict Sense (kTSS) language models (LM), has been proposed in previous works to be integrated in Continuous Speech Recognition (CSR) Systems. In this work, a pruning procedure was applied to k-TSS models in order to reduce the size of the model while keeping its accuracy. An experimental evaluation of the pruned k-TSS models was carried out over a Spanish speech corpus in terms of both, perplexity and word recognition rates. Several pruning thresholds as well as several values of the well-known empirical scaling factor applied to the estimated LM probabilities were tested. These experiments showed that pruned models achieved a good system performance. They also showed that an increase of the test set perplexity of a language model does not always mean a degradation in the model performance when integrated in a CSR system.