Preliminary Evaluation of Angry Voice in Automatic Speech Recognition

Rikio Ueno, Mitsunori Mizumachi, Yoshihisa Nakatoh · 2013

Speech recognition has been introduced as an interface for the various devices; especially operator assistance in call center operations is needed. But when speech recognition is introduced into the call center operations, the recognition performance may deteriorate because the voices of customers include emotion (angry). Previous study reported that the recognition performance of “angry voice” tend to deteriorate than that of “calm voice”. The acoustic features of “angry voice” are different from those of “calm voice”, for example, loud power and high voice. In this study, to explore what factors make the recognition performance deteriorating, we record the parallel speech corpus of “calm voice” and “angry voice” in Japanese, carry out the recognition experiments. And, we compare speech pitch, speech power and spectral envelope between speaker of little deteriorating of speech recognition rate and speaker of deteriorating of speech recognition rate for five vowels. In the results, about speaker deteriorating recognition rate, speech pitch of /i/ increased (about 5dB) and speech power of /u/ increased (about 40Hz) on “angry voice” of “incorrectly words”. Particularly, it was confirmed that the spectral envelopes of /i/ and /u/ on “angry voice” were changed the form.

Read the paper · More papers on PaperTik