Features for speaker-independent recognition of noisy and Lombard speech
Ted H. Applebaum, Brian A. Hanson · The Journal of the Acoustical Society of America · 1990
Additive noise and noise-induced changes in vocal effort (Lombard effect) cause significant loss of performance for recognizers trained on “normal” noise-free speech when the speech is represented by cepstral coefficients or the combination of cepstral coefficients and their first time derivative. The goal of this work is to find a representation of speech that is more robust to the mismatch of test and training noise conditions. In Applebaum and Hanson [EUSIPCO-90, Barcelona], it was shown, for a digits vocabulary, that adding a second derivative of the cepstral coefficients substantially improved recognition rate for a recognizer utilizing perceptually based LP analysis when the derivatives were calculated over long (> 200 ms) regression windows. It was conjectured that these long windows not work well for a more confusable vocabulary. The current study extends this previous work to a 21-word vocabulary consisting of confusable subsets of the English alpha-digits and the words “no” and “go.” Despite the small duration of the discriminating portion of many of the contrasts involved in this vocabulary, the second derivative feature is again found to yield substantial improvement to recognition rate at long regression window lengths. Unlike the previous (digit vocabulary) study, incorporating the third derivative is found to give further improvement. Results are examined separately by noise condition and confusable subset.