An Incremental Selection Method for Semi-Supervised Speaker Adaptation in Speech Emotion Recognition

Carlos M. Castorena, Máximo Cobos, Francesc J. Ferri · IEEE Signal Processing Letters · 2025

Adapting Speech Emotion Recognition (SER) to new, previously unseen speakers, remains a significant challenge due to the variability in emotional expression across speakers and the scarcity of labeled data for adaptation. This study introduces a novel incremental adaptation framework designed to address these challenges by leveraging a modified k-means algorithm to iteratively select representative samples in a latent space. By progressively refining the model in this way, the method facilitates the alignment of a given source domain to a new speaker enhancing generalization. Experiments conducted using data from diverse datasets, under both balanced and unbalanced conditions, demonstrate that the proposed approach outperforms random selection and labeling, achieving comparable or superior results to state-of-the-art, non-incremental methods. These findings underscore the potential of the proposed incremental strategy for improving speaker adaptation in SER tasks, particularly in data-limited scenarios.

Read the paper · More papers on PaperTik