Deep Learning-Based Speech Emotion Recognition: Evaluating CNN, CLSTM, and LSTM Models

Chaitanya Jannu, D. Bharath, V. Sanjay, Himani Kothakotamaana, Srimannarayana Adapa, Veeraswamy Parisae · 2025

Speech emotion recognition (SER) is a critical task in human-computer interaction, with applications ranging from healthcare to customer service. However, accurately identifying emotions from speech signals remains a challenging problem due to the complexity and variability of emotional expressions. To address this challenge, this study proposes a deep learning-based approach for automatic speech emotion recognition using Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). The proposed model employs a CNN, LSTM, and CLSTM models to capture both spatial and temporal features from speech signals, enhancing the accuracy and robustness of emotion classification. While both CNN and RNNs are shown to be effective at obtaining raw audio waveform representations, we examined CNN and RNN unified as a framework in the hope of enhancing their performance characteristics. The CNN is trained to learn detailed, low-level speech representations directly from raw waveform data, without through pre-engineered features or through spectral representations. This allows for the capture of narrow-band characteristics that are indicative of emotion. However, while CNN components extract detailed time frames of data, the LSTM layers themselves model temporal dynamics. The study leverages a combined dataset comprising the TESS, RAVDESS and Combined datasets, which include a diverse range of emotional speech samples such as neutral, happy, sad, angry, fearful, disgust, and surprise. Tested with the TESS dataset, the proposed model surpassed both classification schemas and state-of-the-art baselines, with experiments highlighted by the accuracy of emotion recognition.

Read the paper · More papers on PaperTik