Incorporation of Temporal Structure Into a Vector-Quantization-Based Preprocessor for Speaker-Independent, Isolated-Word Recognition
A. Bergh, Frank K. Soong, L. R. Rabiner · AT&T Technical Journal · 1985
Recently a new structure for isolated-word recognition was proposed in which a separate Vector Quantization (VQ) code book was designed for each word in the vocabulary. The word-based VQs were used as a front-end preprocessor to eliminate word candidates whose distortion scores were large; a dynamic time-warping processor then resolved the choice among the remaining word candidates. The above scheme worked very well for small vocabularies; however, the major flaw was the lack of temporal information in the word-based VQ processor. As such, as the vocabulary grew in size and complexity, the ability of the VQ processor to resolve among similar sounding words decreased dramatically, and the effectiveness of the proposed recognition structure similarly decreased. To alleviate this difficulty a technique for incorporating temporal structure into the preprocessor is proposed. In particular, the probability density function of the time of occurrence for each vector in the code book is estimated from a training sequence. In the recognizer, the spectral distance score of the VQ is combined with a temporal distance score, for each frame in the word. An evaluation of the modified recognizer showed slightly improved performance on the digits vocabulary and greatly improved performance on a vocabulary of 129 airlines terms.