Using the Fisher Vector Representation for Audio-based Emotion Recognition
Gábor Gosztolya · Acta Polytechnica Hungarica · 2020
Automatically determining speaker emotions in human speech is a frequently studied task, where various techniques have been employed over the years.An efficient method is to represent the utterances by employing the Bag-of-Audio-Words technique, inspired by the Bag-of-Visual-Words approach from the area of image processing.In the past few years, however, Bag-of-Visual-Words has been replaced by the so-called Fisher vector representation, as it was shown to give a better classification performance.Despite this, in audio processing, Fisher vectors to date have only been rarely applied.In this study, we show that Fisher vectors are also a viable way of representing features in speech technology; more precisely, we use them in the task of emotion classification.Based on our results on two datasets, Fisher vectors can be effectively employed for this task: we measured 4% relative improvements in the UAR scores for both corpora, which rose to 9-16% when we combined this approach with the standard paralinguistic one.