K-TLSS(S) language models for speech recognition

Germán Bordel, Amparo Varona, M. Inés Torres · 2002

The class of K-testable languages in the strict sense (K-TLSS) is a subclass of the regular languages. Stochastic K-TLSS language models describe the same probability distribution as N-gram models, and smoothing techniques (backoff-like methods) can be applied efficiently. Once we have a set of k-TLSS models (k=1...K) and a smoothing technique that specifically fits them, we propose an integration into a unique self-contained [K-TLSS(S)] model which embeds the smoothing within the topology, allowing extremely simple parsing procedures. To build this model, we designed a more general syntactic mechanism that we call a "stochastic deterministic finite state automaton with recursive transitions". The topology of the new K-TLSS(S) model allows an easy pruning procedure. Pruned K-TLSS(S) models give probability distributions that are equivalent to variable-length N-gram models. Experimental results give as a conclusion that the effect of a small pruning is always positive.

Read the paper · More papers on PaperTik