A Study of What Makes Calm and Sad Music So Difficult to Distinguish in Music Emotion Recognition.

Yu Hong, Chuck-jee Chau, Andrew Horner · Rare & Special e-Zone (The Hong Kong University of Science and Technology) · 2017

Music emotion recognition and recommendation systems often use a simplified 4-quadrant model with categories such as Happy, Sad, Angry, and Calm. Previous research has shown that both listeners and automated systems often have difficulty distinguishing low-arousal categories such as Calm and Sad. This paper seeks to explore what makes the categories Calm and Sad so difficult to distinguish. We used 300 low-arousal excerpts from the classical piano repertoire to determine the coverage of the categories Calm and Sad in the low-arousal space, their overlap, and their balance to one another. Our results show that Calm was 50% bigger in terms of coverage than Sad, but that on average Sad excerpts were significantly more negative in mood than Calm excerpts were positive. Calm and Sad overlapped in nearly 20% of the excerpts, meaning 20% of the excerpts were about equally Calm and Sad. Calm and Sad covered about 92% of the low-arousal space, where 8% of the space were holes that were not-at-all Calm or Sad. Due to the holes in the coverage, the overlaps, and imbalances, the Calm-Sad model adds about 4% more errors when compared to asking users directly whether the mood of the music is positive or negative.

Read the paper · More papers on PaperTik