Application of prosody modification for Speech Recognition in different Emotion conditions

Vishnu Vidyadhara Raju, Paidi Gangamohan, Suryakanth V. Gangashetty, Anil Kumar Vuppala · 2016

The main focus of this study is to analyze the performance of Automatic Speech Recognition (ASR) in different emotional environments using prosody modification. The majority of ASR systems are trained using neutral speech and the performance of such systems degrade when tested with the emotional speech. In this paper, the various components of speech that contribute to the emotion characteristics are studied. The prosody features of the source emotional utterances are modified according to the target neutral utterances using Flexible Analysis Synthesis Tool (FAST). In the FAST, Dynamic Time Warping (DTW) is used to align the source emotional and target neutral utterances. Components of the prosody such as intonation, duration and excitation source are manipulated to incorporate the desired features into the source utterance. The modified (source emotional) utterances are then used for testing the ASR system which is trained using neutral speech. In this study, three emotions (compassion, happiness and anger) are considered for the analysis. Experimental results indicate an average improvement in the speech recognition system performance by considering prosody modified speech.

Read the paper · More papers on PaperTik