Acoustic recognition component of an 86000-word speech recognizer
Li Deng, V. Gupta, Matthew Lennig, Patrick J Kenny, Paul G. Mermelstein · International Conference on Acoustics, Speech, and Signal Processing · 2002
Recent results obtained with a hidden Markov model (HMM)-based acoustic recognizer using a virtually unlimited vocabulary (86000 words) to perform speaker-dependent isolated-word recognition are described. The task domain of this recognizer is quite general, consisting of paragraphs read from various newspapers, books, and magazines. The results of a comparative acoustic recognition study using various types of HMMs and various amounts of training data (from 700 to about 4000 words) are presented. The models explored include context-dependent allophonic HMMs (including generalized diphone and triphone models with unimodal Gaussian output densities) and context-independent phonemic HMMs (using either unimodal or mixture densities). Experimental results indicate that phonemic HMMs with many components in the mixture output densities provide the highest acoustic recognition accuracy. The acoustic recognition accuracy for a total of about 7000 test words spoken by four male and five female speakers is 82%. Recognition accuracy after application of the language model increases to 92%.>