Speaker Normalisation for Speech-Based Emotion Detection
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps · 2007
The focus of this paper is on speech-based emotion detection utilising only acoustic data, i.e. without using any linguistic or semantic information. However, this approach in general suffers from the fact that acoustic data is speaker-dependent, and can result in inefficient estimation of the statistics modelled by classifiers such as hidden Markov models (HMMs) and Gaussian mixture models (GMMs). We propose the use of speaker-specific feature warping as a means of normalising acoustic features to overcome the problem of speaker dependency. In this paper we compare the performance of a system that uses feature warping to one that does not. The back-end employs an HMM-based classifier that captures the temporal variations of the feature vectors by modelling them as transitions between different states. Evaluations conducted on the LDC emotional prosody speech corpus reveal a relative increase in classification accuracy of up to 20%.