Transform representation of the spectra of acoustic speech segments with applications. I. General approach and application to speech recognition
V. Ralph Algazi, Kathy L. Brown, Michael J. Ready, David H. Irvine, C.L. Cadwell, Sanguoon Chung · IEEE Transactions on Speech and Audio Processing · 1993
An approach to modeling and capturing the time-varying structure of the spectral envelope of speech is reported. Acoustic subword decomposition and the Karhunen-Loeve transform (KLT) are used to extract and efficiently represent the highly correlated structure of the spectral envelope. Integration of the KLT with acoustic subword modeling provides concise representation of both steady-state and dynamic features of the spectra in a unified framework that very effectively captures acoustic-phonetic patterns. The physiological and perceptual basis for the approach, the frame-based and acoustic-subword-based spectral representation, and applications to speaker-dependent recognition are presented. The performance of the recognition algorithm based on this approach compares favorably with that of other techniques.>