Investigations on search methods for speech recognition using weighted finite state transducers
David Rybach, Hermann Ney · RWTH Publications (RWTH Aachen) · 2014
The search problem in the statistical approach to speech recognition is to find the most likely word sequence for an observed speech signal using a combination of knowledge sources, i.e. the language model, the pronunciation model, and the acoustic models of phones. The resulting search space is enormous. Therefore, an efficient search strategy is required to compute the result with a feasible amount of time and memory. The structured statistical models as well as their combination, the search network, can be represented as weighted finite-state transducers. The combination of the individual transducers and their optimization for size and structure is achieved by means of weighted transducer algorithms. The construction of the transducer can be performed on-the-fly during the search, such that only parts of the search network are generated as needed. This dynamic network search has lower memory requirements compared to a search using the full static expansion of the search network. In this thesis, we investigate search methods for speech recognition using weighted finite-state transducers. The focus of this work is on dynamic search networks using on-the-fly transducer composition. We study the construction of the transducers involved and analyze different modeling approaches. Amongst other topics, we describe a novel construction of compact phone context-dependency transducers based on a joint optimization of model complexity and transducer size. We describe an efficient search algorithm and its implementation in detail and provide an experimental evaluation. The dynamic transducer-based search is compared in-depth to another state-of-the-art search strategy using dynamic network expansion, namely history conditioned lexical tree search. Experimental results are obtained using several high-performance large vocabulary continuous speech recognition systems, including systems for broadcast news and spontaneous speech in English and Arabic. This thesis includes also considerations on practical aspects of a speech recognition system. In particular, we describe a novel framework for audio segmentation and we give a detailed overview of RASR, the publicly available RWTH Aachen University speech recognition software package, which has been extended within the scope of this work.