Extracting, modelling and combining information in speech recognition

Kris Demuynck · Lirias · 2001

The task of a speech recogniser is to transcribe human speech into text. To do so, modern recognisers rely firmly on the principles of statistical pattern recognition. This statistical framework allows the problem of speech recognition to be decomposed into a set of well-de ned sub-tasks, namely the extraction of the relevant features from the acoustic signal, the statistical modelling of the acoustic and linguistic knowledge, and the combination of all information in an efficient search for the most probable sentence. As is indicated in the title, this dissertation addresses topics from all sub-domains. The first objective of our work is the creation of a basic recognition system suitable for advanced research in the domain of speaker independent large vocabulary continuous speech recognition. Being a research system, flexibility and ease of use (few tuning parameters) are the primary design criteria. However, since research in the domain of speech recognition relies heavily on experimental validation, the decoding speed and the memory efficiency of the system are of importance as well. The two most critical building blocks are the acoustic models and the search-engine (decoder). For the acoustic models, we opted for "reduced semi-continuous" hidden Markov models (HMMs) since these offer the required flexibility and hold the promise to be memory efficient. Our work concerning this type of modelling consists of two main parts. In the first part, two methods to construct the initial set of gaussians as needed to create "reduced semi-continuous" HMMs, are presented. In the second part, a method to accelerate the evaluation of these acoustic models with a factor of 10 and more, is given. For the decoder, a novel search-topology is presented which consist of a memory efficient network structure to store the pronunciation information (lexicon in terms of context dependent units) in combination with a run-time (dynamic) integration of the language model information. This topology offers fast access to all relevant data but has nevertheless only modest memory requirements. The topology also imposes hardly any constraints on the different knowledge sources. The resulting system is capable of handling all available knowledge sources (intra- and crossword context dependent phone models, long-span language models and assimilation rules) in a single time-synchronous recognition pass, and this at a competitive decoding speed. The validity of all design options (e.g. the use of "reduced semicontinuous" HMMs and the topology of the decoder) is verified by comparing our recogniser with other state-of-the-art systems on some standardised benchmark test suites. The second part of the work concerns some advanced research on the interaction between preprocessing and acoustic modelling in order to further enhance the accuracy of the recognition system developed so far. As a result of this research, two linear transformations were added to the preprocessing scheme. The first transformation is an alternative to the commonly used linear discriminant analysis (LDA) for finding a compact and discriminative set of features. The main advantage of the new scheme over LDA is the use of a more powerful criterion, namely a minimal loss in information content of the reduced feature set with respect to the original feature set. The second transformation compensates for the limited capabilities of the acoustic models when it comes to handling correlated data. Each of these transformations improves the recognition accuracy with 5 to 10% relative while adding less than 0.5% parameters to the acoustic modelling.

Read the paper · More papers on PaperTik