Long-term spectro-temporal information for improved automatic speech emotion classification

Siqing Wu, Tiago Henrique Falk, Wai-Yip Chan · 2008

This paper investigates the contribution of features which con-vey long-term spectro-temporal (ST) information for the pur-pose of automatic emotional speech classification. The ST rep-resentation is obtained by means of a modulation filterbank de-composition of long-term temporal envelopes of the outputs of a gammatone filterbank. The two-dimensional discrete cosine transform is used to reduce the dimensionality of the represen-tation; candidate features are then derived from statistics com-puted from the DCT coefficients. Sequential forward feature selection is used to select the most salient features. Two types of experiments are described which use the Berlin emotional speech database to test the performance of the ST features alone and in combination with prosodic features. In a multi-class experiment, simulation results with a support vector classifier show that a 44 % reduction in classification error is attained once prosodic features are combined with the proposed ST fea-tures. Additionally, in a one-against-all experiment, an average increase in F-score of 33 % is attained when the proposed ST features are included. Index Terms: speech emotion recognition, spectro-temporal features, modulation spectrum, affective computing.

Read the paper · More papers on PaperTik