The Munich LSTM-RNN Approach to the MediaEval 2014 "Emotion in Music" Task
Eduardo Coutinho, Felix Johannes Weninger, Björn Wolfgang Schuller, Klaus R. Scherer · 2014
In this paper we describe TUM’s approach for the Medi-aEval’s “Emotion in Music ” task. The goal of this task is to automatically estimate the emotions expressed by mu-sic (in terms of Arousal and Valence) in a time-continuous fashion. Our system consists of Long-Short Term Mem-ory Recurrent Neural Networks (LSTM-RNN) for dynamic Arousal and Valence regression. We used two different sets of acoustic and psychoacoustic features that have been previ-ously proven as effective for emotion prediction in music and speech. The best model yielded an average Pearson’s cor-relation coefficient of 0.354 (Arousal) and 0.198 (Valence), and an average Root Mean Squared Error of 0.102 (Arousal) and 0.079 (Valence).