Research on speech emotion recognition based on deep auto-encoder
Wang Fei, Xiaofeng Ye, Zhao‐Yu Sun, Yujia Huang, Xing Zhang, Shengxing Shang · 2016
Good features are critical for the research of speech emotion recognition. This paper based on the theory of deep learning, and phonetic features were extracted by using the method of deep auto-encoder (DAE). In this paper, a deep auto-encoder containing five hidden layers was designed. To get the input data, we divided the audio into short frames, each frame of speech emotion signal was then decomposed with wavelet, and was calculated the Fourier transform. Higher features were learned with deep automatic encoder, and some traditional features such as MFCC, LPCC were also extracted. With high-level features and traditional features, support vector machine was used for classification and recognition. Compared with the traditional features, the results show that the highest recognition accuracy rate can be reached 86.41%.