Voice Quality Dependent Speech Recognition
Tae Jin Yoon, Xiaodan Zhuang, Jennifer Cole, Mark Hasegawa‐Johnson · 2009
Voice quality conveys both linguistic and paralinguistic information, and can be distinguished by acoustic source characteristics. We label objective voice quality categories based on the spectral and temporal structure of speech sounds, specifically the harmonic structure (H1-H2) and the mean autocorrelation ratio of each phone. Results from a classification experiment using a Support Vector Machine (SVM) show that allophones that differ from each other regarding voice quality can be classified using input features in speech recognition. Among different possible ways to incorporate voice quality information in speech recognition, we demonstrate that by explicitly modeling voice quality variance in the acoustic models using hidden Markov models, we can improve word recognition accuracy. Keywords: ASR, Voice quality, H1-H2, Autocorrelation ratio, SVM, HMM. 1