Robust Acoustic Speech Emotion Recognition by Ensembles of Classifiers

Björn Wolfgang Schuller, M. Lang, Gerhard Rigoll · OPUS (Augsburg University) · 2005

Automatic speech recognition can fail to a certain extent when confronted with emotionally distorted speech. Great efforts have been spent so far to cope with noise conditions or speaker’s characteristics. Yet, adaptation to the emotional condition of the speaker could help to further improve the overall performance. In this respect we aim at a robust and reliable recognition of the speaker’s emotional state by acoustic features only prior to speech recognition itself. Thereby we can load according emotional speech models. In this work we introduce an optimal feature set for this task selected by Sequential Floating Search Methods. The set comprises high-level prosodic features resembling utterancewise statistic analysis of low-level contours as pitch, higherorder formants, energy, and spectral development. Within classification we apply ensemble classification as Stacking, Bagging, and Boosting.

Read the paper · More papers on PaperTik