ON-LINE TRANSDUCER COMPOSITION AND SMOOTHED LANGUAGE MODEL INCORPORATION
Daniel Willett, Shigeru Katagiri · 2002
This paper presents and evaluates our recent efforts on ef ficient decoding for Large Vocabulary Continuous Speech Recognit ion in the framework of Weighted Finite State Transducers. We evaluat e on-the-fty transducer compo sition for reduced memory consumption combined with weight smearing for a more time-synchron ous language model incorporation. It turns out that in the on-li ne com position mode weight smoothing within the static part of the network is even more beneficial on run-time to accu racy ratio than in the fully precompiled case. Evaluations are carried out on a state-of-the-art recognition system of 10k words, cross-word triphone acoustie models and trigram language model. In this scenario, the Viterbi-search is car ried out full y time-synchro nously in only a single pass. The combination of on-the-fly network composition with only the unigram part of the language model smoothly compiled into the network achieves a remarkably good run-time to accuracy ratio with only moderate memory requirements.