On the use of triphone models for continuous speech recognition
Kai-Fu Lee · The Journal of the Acoustical Society of America · 1988
This paper describes both completed and future work with triphone models for speech recognition. Triphone models are a powerful subword modeling technique because they account for the left and right phonetic contexts. One problem with triphone models is that there are too many models for the amount of training data available. Moreover, many triphone contexts are very similar. In view of this, generalized triphones are introduced, which are created by clustering similar triphones together using an information theoretic criterion. This technique provides a way of finding the “right” number of models for any task. With generalized triphone models, a 96% accuracy with a 1000-word speaker-independent continuous task was achieved. There are only 2500 triphones for the above task. For a different task, there will be new triphones; for a natural English task, the number of triphones will be much larger. In the future, improvement of the applicability of triphone modeling could be done by (1) training a large set of generalized triphone models for English, and (2) exploring techniques to deal with new triphones. [Work supported by DARPA.]