Empirical mode decomposition based weighted frequency feature for speech-based emotion classification

Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps · IEEE International Conference on Acoustics Speech and Signal Processing · 2008

This paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like rate of speech. We propose the use of a novel feature based on instantaneous frequency obtained from the speech, in addition to the aforementioned features, in order to take into account the vocal tract parameters as well as vocal chord excitation. The proposed features employ the recently emerged empirical mode decomposition to decompose speech into AM-FM signals that are symmetric about zero and suitable for Hilbert transformation to extract the instantaneous frequency. The proposed features provide a relative increase in classification accuracy of approximately 9% when appended to established acoustic features.

Read the paper · More papers on PaperTik