Emotion Recognition in Speech with Latent Discriminative Representations Learning

Jing Han, Zixing Zhang, Gil Keren, Björn Wolfgang Schuller · Acta acustica united with Acustica · 2018

Despite significant recent advances in the field of affective computing, learning meaningful representations for emotion recognition remains quite challenging.In this paper,wepropose anovelfeature learning approach named Latent Discriminative Representation (LDR)learning for speech emotion recognition.Unlikemost existing handcrafted features designed for specificapplications or features learnt by astandard neural network, the proposed learning method incorporates an additional training objective in order to learn better representations of the task of interest.To this end, we group the training samples into sets of triplets, satisfying that the second member in each triplet comes from the same class as the first and that the third member comes from ad if ferent class than the first.In the training pr ocess, we maximise the distance of the samples from different classes in the latent representation space, while we minimise the distance for samples from the same class.To evaluate the effectiveness of LDR, we perform extensive experiments on the widely used database IEMOCAP,and findthat the LDR improvesperformance overthe standard neural network training procedure.

Read the paper · More papers on PaperTik