Towards a Corpus of Speech Emotion for Interactive Dialog Systems
Dario Bertero, Farhad Bin Siddique, Pascale Fung · 2017
We present and discuss an ongoing data collection and annotation effort to build a large corpus on speech emotion detection. We collected 207 hours of public speech data from TED talks. We highlight the expected relation between the API output emotion distribution and the common features of a TED talk. We then employed manual annotators to improve the quality of the annotations, building a two-level annotation process where to a main emotion label we add multiple secondary emotion descriptors. Additional annotations were also obtained through crowdsourcing, and we provide an analysis of the agreement between multiple annotators. We conducted a comparison between the automatic and manual labeling, obtaining an average accuracy of 28.4% of the automatic multiclass annotation API, as well as automatic classification with a CNN, reaching more than 60% average accuracy in various experimental settings. We also discuss various active learning methods we are using to select the samples to be annotated in order to obtain more relevant data at a faster pace. A large speech emotion detection corpus will enable more accurate emotion detection systems, which can then be integrated into dialog systems to recognize and react to the user emotion.