HMM based acoustic modelling in large vocabulary speech recognition
Jacques Duchateau · Lirias · 1998
In general the aim of an automatic speech recognition system is to write down what is said. State of the art continuous speech recognition systems for large vocabulary - for an entire language - consist of four basic modules: the signal processing, the acoustic modelling, the language modelling and the search engine. The subject of this thesis is the acoustic modelling module, which models the different sounds (phonemes) in the language. It is investigated how the model design strategy and the algorithms can be adapted to a specific type of models called Semi-Continuous Hidden Markov Models (HMMs)}. This type of modelling is more general and more flexible than what is often used in nowadays recognisers. Important topics in acoustic modelling are the accuracy of the models and their evaluation speed. As for the evaluation speed, the so called FRoG algorithm and reduced Semi-Continuous HMMs are proposed to make the evaluation of the models considerably faster. For the accurate acoustic modelling, a strategy is adopted with cross-word context dependent phoneme models based on phonetic decision trees. The strategy and algorithms are adapted whenever necessary to the modelling with Semi-Continuous HMMs, among other things proposing a new algorithm for the node splitting criterion used during decision tree construction. The proposed algorithms are also compared with more standard techniques, evaluating on benchmark recognition tests for large vocabulary recognition as used in the speech recognition community. Although a lot of research is carried out for English - as the benchmark recognition tasks are for English - special attention is also paid to large vocabulary recognition for Dutch, for instance cooperating in the development of linguistic resources that are necessary to design a recogniser for Dutch.