Speech Emotion Recognition Using Semi-Supervised Learning with Efficient Labeling Strategies
Zhi Zhu, Yoshinao Sato · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) · 2021
The collection of large amounts of labeled data for speech emotion recognition requires considerable time and effort. As a result, the sizes of existing corpora are limited. One promising solution to this difficulty is semi-supervised learning, i.e., learning from both labeled and unlabeled data. In this study, we applied the noisy student training (NST) method to speech emotion recognition. We experimentally investigate the trade-off between the amount and reliability of labeled data. For this purpose, we prepared labeled and unlabeled data by lim-iting the available annotations in the CREMA-D dataset. The experimental results showed that a model trained using the NST method with some of the annotations achieved almost the same performance as the one trained using supervised learning with all the annotations if the amount and reliabil-ity of the available annotations were appropriate. Our findings are significant in identifying the most efficient labeling strategy when utilizing a large-scale dataset without labels for speech emotion recognition.