Recognition of emotional speech with convolutional neural networks by means of spectral estimates

Norman Weiskirchen, Ronald Böck, Andreas Wendemuth · 2017

Current developments in deep neural architectures achieved remarkable results in the classification of emotions from speech. Recently, also cross-modal approaches gained attention in the community. Such a classification method is the Convolutional Neural Network (CNN). Mainly developed for analyses of images it can be used also in speech processing. In this paper, we present a CNN-based classification architecture adapting spectrograms as representations of emotion-afflicted speech input. Given this approach, we applied our network architecture to three benchmark corpora, namely EmoDB, eNTERFACE, and SUSAS, and investigated the classification ability in a Leave-One-Speaker-Out setting. Especially, for SUSAS, a close-to-real-life corpus, remarkable results were obtained. In addition, we investigated the option of analysing CNN's internal representations of the given input using Deep Dreaming. For this, we were able to identify spectral parts which contribute most to the classification process.

Read the paper · More papers on PaperTik