A hybrid segmental neural net/hidden Markov model system for continuous speech recognition
George Zavaliagkos, Ying Zhao, Richard M. Schwartz, J. Makhoul · IEEE Transactions on Speech and Audio Processing · 1994
The current state-of-the-art in large-vocabulary, continuous speech recognition is based on the use of hidden Markov models (HMM). In an attempt to improve over HMM performance, the authors developed a hybrid system that combines the advantages of neural networks and HMM using a multiple hypothesis (or N-best) paradigm. The connectionist component of the system, the segmental neural net (SNN), models all the frames of a phonetic segment simultaneously, thus overcoming the well-known conditional-independence limitation of the HMM. They describe the hybrid system and discuss various aspects of SNN modeling, including network architectures, training algorithms and context modeling. Finally, they evaluate the hybrid system by performing several speaker-independent experiments with the DARPA Resource Management (RM) corpus, and demonstrate that the hybrid system shows a consistent improvement in performance over the baseline HMM system.