Speech recognition using least-squares models for the contextual variations in the acoustic characteristics of phonetic segments

Yoshiharu Abe, Kunio Nakajima · The Journal of the Acoustical Society of America · 1988

This paper describes a speech recognition method using linear models to predict the variations in a feature vector caused by the phonemic context. The feature vector in a phonetic segment is decomposed and formulated by the sum of a context-independent vector, a context-dependent vector, and an estimation error. The second component is given by the product of a weight matrix and an environment vector. Least-squares estimations of the context independent vectors and the weight matrices for top-down, bottom-up, and mixed models are obtained algebraically. The top-down model obtains the environment vector from the phoneme string transcribed in the lexicon, while the bottom-up model obtains it from the input feature vectors. Speaker-dependent recognition experiments were carried out using an open vocabulary of 200 words spoken by two male and two female speakers. The average word error rate of 19.1% without any contextual models was reduced to 7.4%, 5.3%, and 3.3% by the top-down, bottom-up, and mixed models, respectively.

Read the paper · More papers on PaperTik