An improved HMM/VQ training procedure for speaker-independent isolated word recognition
Yaxin Zhang, Michael D. Alder · 2002
This paper describe an improved training procedure in a HMM/VQ speech recognition system for speaker-independent speech recognition. The phoneme based Gaussian mixture models (GMM) were generated in the first step modeling using the Expectation-Maximization (EM) algorithm. These Gaussians more accurately describe the distribution characteristic of the phonemes in the speech signal space. Therefore better first step modeling is achieved and the performance of the whole recognition system is improved. The new method was used in a speaker-independent isolated digits and phoneme recognition tasks. Two English databases were used for the training and testing. Significant improvements have been achieved in comparison with the conventional HMM/VQ system.>