A Comprehensive Analysis of Data Augmentation Methods for Speech Emotion Recognition

Umut Avcı · IEEE Access · 2025

The limited availability of labeled emotional speech data remains a significant challenge in developing robust speech emotion recognition systems. This paper presents a comprehensive investigation into the effectiveness of diverse data augmentation strategies for enhancing emotion recognition performance. Three different data augmentation categories are examined: audio-based transformations, image-based modifications, and feature-level synthesis. Seventeen transformations are used in audio-based data augmentation to change the raw audio signal’s time and frequency content. Eight transformations, such as shifting, rotating, and zooming, are applied to the spectrogram images for image-based data augmentation. The SpecAugment method is also used to transform the spectrograms into versions with masked time and frequency axes. In feature-space-based approaches, new feature vectors are generated using five oversampling algorithms and a generative adversarial network. Experimental results from EMO-DB and IEMOCAP datasets demonstrate that the data augmentation approaches enhance emotion classification performance by up to six percent. The empirical evidence indicates that training sets augmented through combinations of audio-based transformations yield the highest performance gains. In contrast, the GAN-based approach fails to produce improvements in classification performance.

Read the paper · More papers on PaperTik