Efficient Generation of high-order context-dependent Weighted Finite State Transducers for Speech Recognition

Maria Elke Schuster, Takaaki Hori · 2006

This paper describes an algorithm for efficient building of weighted finite state transducers for speech recognition when high-order context-dependent models of order K>3 (triphones) with tied states are used. We show how an algorithm to build a part of the needed composed transducers directly from the decision trees in combination with an improved compilation process can lead to much faster, simpler and more memory-efficient compilation. In our case, it also resulted in substantially smaller final networks. With the described algorithm, it is simple to use high-order full cross-word models with little overhead directly within a one-pass time-synchronous search, which we test comparing resulting final network sizes, recognition rates and speed on a large, spontaneous Japanese speech database. Using the proposed algorithm, it is possible to do real-time recognition using full crossword quinphones with a large acoustic model in about 125 MB of memory at about 9% search error.

Read the paper · More papers on PaperTik