Emotional speech classification with prosodic prameters by using neural networks

Hirokazu Sato, Yasue Mitsukura, Minoru Fukumi, Norio Akamatsu · 2001

Interestingly, in order to achieve a new Human Interface such that digital computers can deal with the KASEI information, the study of the KANSEI information processing recently has been approached. In this paper, we propose a new classification method of emotional speech by analyzing feature parameters obtained from the emotional speech and by learning them using neural networks, which is regarded as a KANSEI information processing. In the present research, KANSEI information is usually human emotion. The emotion is classified broadly into four patterns such as neutral, anger, sad and joy. The pitch as one of feature parameters governs voice modulation, and can be sensitive to change of emotion. The pitch is extracted from each emotional speech by the cepstrum method. Input values of neural networks (NNs) are then emotional pitch patterns, which are time-varying. It is shown that NNs can achieve classification of emotion by learning each emotional pitch pattern by means of computer simulations.

Read the paper · More papers on PaperTik