Pre-initialized composition for large-vocabulary speech recognition

Cyril Allauzen, Michael Riley · 2013

This paper describes a modified composition algorithm that is used for combining two finite-state transducers, representing the context-dependent lexicon and the language model respec-tively, in large vocabulary speech recogntion. This algorithm is a hybrid between the static and dynamic expansion of the re-sultant transducer, which maps from context-dependent phones to words and is searched during decoding. The approach is to pre-compute part of the recognition transducer and leave the balance to be expanded during decoding. This method allows for a fine-grained trade-off between space and time in recogni-tion. For example, the time overhead of purely dynamic expan-sion can be reduced by over six-fold with only a 20 % increase in memory in a collection of large-vocabulary recognition tasks available on the Google Android platform.

Read the paper · More papers on PaperTik