Emotions in Speech - Experiments with Prosody and Quality Features in Speech for Use in Categorical and Dimensional Emotion Recognition Environments

Mark Borchert, Antje Düsterhöft · 2006

This paper focuses on features in speech and classification algorithms for using them in emotion recognition software. We illustrate a new approach concentrated on analyzing speech quality features. The quality features are formants, spectral energy distribution in different frequency bands, harmonics-to-noise ratio (in different frequency bands) and irregularities (jitter, shimmer). Some papers (A. Dusterhoft et al., 2003, D. Goleman, 1995) show that there is a relationship between quality features and the valence axis. This paper deals with dimensional approach to classify emotions. Therefore, mainly quality features are taken for the valence axis to classify emotions and mainly prosody features are taken for the arousal axis. Because our experiments show that single emotion recognition rates are up to 90 percent and the recognition rates for speaker independent recognition is about 70% for all classification algorithm, it seems that quality features are more appropriate to differentiate emotions with the same arousal and different valence levels in a dimensional approach. A prototypical emotion recognition software is implemented which is actually tested for analyzing the mood of customers in call centers.

Read the paper · More papers on PaperTik