High-arousal emotional speech enhances speech intelligibility and emotion recognition in noise
Jessica M. Alexander, Fernando Llanos · The Journal of the Acoustical Society of America · 2025
Prosodic and voice quality modulations of the speech signal offer acoustic cues to the emotional state of the speaker. In quiet, listeners are highly adept at identifying not only a speaker's words but also the underlying emotional context. Given that distinct vocal emotions possess varying acoustic characteristics, background noise level may differentially impact speech recognition, emotion recognition, or their interaction. To investigate this question, we assessed the effects of three emotional speech styles (angry, happy, neutral) on speech intelligibility and emotion recognition across four different SNR levels. High-arousal emotional speech styles (happy and angry speech) enhanced both speech intelligibility and emotion recognition in noise. However, emotion recognition behavior was not a reliable predictor of speech recognition behavior. Instead, we found a strong correspondence between speech recognition scores and the relative power of the speech-in-noise signal in critical bands derived from the Speech Intelligibility Index. Unsupervised dimensional scaling analysis of emotion recognition patterns revealed that different noise baselines elicit different perceptual cue weighting strategies. Further dimensional scaling analysis revealed that emotion recognition patterns were best predicted by emotion-level differences in harmonic-to-noise ratio and variability around the fundamental frequency. Listeners may thus weight acoustic features differently for recognizing speech versus emotional patterns.