Dynamic grammars with lookahead composition for WFST-based speech recognition

Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose · 2012

Automatic Speech Recognition (ASR) applications often em-ploy a mixture of static and dynamic grammar components, and can thus benefit from the ability to efficiently modify the sys-tem vocabulary and other parameters in an on-line mode. This paper presents a novel, generic approach to dynamic grammar handling in the context of the Weighted Finite-State Transducer (WFST) paradigm. The method relies on a straightforward ex-tension of the lexicon and underlying grammar components, and leverages the ideas of on-the-fly composition and delayed construction to efficiently generate the recognition search space on-the-fly. The alternative partitioning of component models that this approach implies can also result in significant stor-age savings. In contrast to previous works in this area, the proposed method relies only on generic WFST operations and the context-dependency, lexicon and grammar components that form the basis of standard ASR cascades.

Read the paper · More papers on PaperTik