Predicting Time-Varying Musical Emotion Distributions from Multi-Track Audio

Erik M. Schmidt, Matthew Prockup, Brandon G. Morton, Youngmoo E. Kim · 2012

Music exists primarily as a medium for the expression of emotions, but quantifying such emotional content empirically proves a very dicult task. Myriad features comprise emotion, and as such mu- sic theory provides no rigorous foundation for analysis (e.g. key, mode, tempo, harmony, timbre, and loudness all play some roll), and the weight of individual musical features may vary due to the expressiveness of dif- ferent performers. In previous work, we have shown that the ambiguities of emotions make the determination of a single, unequivocal response label for the mood of a piece of music unrealistic, and we have instead chosen to model human response labels to music in the arousal-valence (A-V) representation of aect as a stochastic distribution. Using multi- track sources, we seek to better understand these distributions by ana- lyzing our content at the performer level for dierent instruments, thus allowing the use of instrument-level features and the ability to isolate af- fect as a result of dierent performers. Following from the time-varying nature of music, we analyze 30-second clips on one-second intervals, in- vestigating several regression techniques for the automatic parameteriza- tion of emotion-space distributions from acoustic data. We compare the results of the individual instruments to the predictions from the entire instrument mixture as well as ensemble methods used to combine the individual regressors from the separate instruments.

Read the paper · More papers on PaperTik