Time and frequency spectral derivative features for robust recognition of Lombard and noisy speech
Brian A. Hanson, Ted H. Applebaum, Gregory R. De Haan · The Journal of the Acoustical Society of America · 1989
Lombard speech approximates speech variations encountered in noisy environments. This paper addresses automatic, speaker-independent recognition of Lombard and noisy speech. Spectral information was derived from production-based (LP) or perceptually based LP analysis, and represented by cepstral or index-weighted cepstral (RPS, spectral frequency derivative) coefficients. Instantaneous, dynamic (first temporal derivative), and acceleration (second temporal derivative) features were then computed from one of the representations, and integrated during the likelihood calculation of a hidden Markov model recognizer. The recognizer was trained with normal speech to evaluate its robustness to speech variations. Strong interaction was found between the temporal derivative features, the spectral derivative, and the degree of smoothing in the analysis. Although the acceleration feature performed poorly by itself, when combined with other features it generally raised recognition rates for Lombard speech. This trend was more pronounced with the cepstral representation. Combining all three features gave the best Lombard speech recognition results: The perceptually based analysis with RPS coefficients and three features yielded 94% correct recognition, compared with 63% and 82% correct for standard cepstrum and RPS, respectively, when the instantaneous feature was used alone. Experiments were also done with additive noise. The results and implications for robust speech recognition are discussed.