Performance Evaluation of Speech Emotion Recognition Using Hybrid Feature Selection and Machine Learning

Ferdinand Mahardhika, Marta Lenah Haryanti, Puguh Hiskiawan · 2025

Speech Emotion Recognition (SER) plays a crucial role in human-computer interaction, offering insights into user emotions, particularly in environments with complex acoustic interactions such as songs. Traditional machine learning classifiers struggle with these complexities, especially when distinguishing emotions in the presence of vocals and background music. This study proposes a hybrid feature selection strategy to enhance SER accuracy for six distinct emotions: neutral, calm, happy, sad, angry, and fearful. Using the RAVDESS database, consisting of 1,012 recordings, we extracted 33 time-frequency features, including energy, spectral descriptors, MFCCs, and chroma. A hybrid framework combining filter-based (chi-square) and embedded (random forest importance) methods was implemented. Four classifiers K Nearest Neighbors (K-NN), Support Vector Machine (SVM), Random Forest (RF), and Gradient Boosting (GB) were evaluated across three paradigms: filter-only, embedded-only, and hybrid. The hybrid feature selection method outperformed the others, with SVM achieving a peak accuracy of 83.55% (RBF kernel) and K-NN showing a 6.25% improvement over baseline performance (81.91%). These results demonstrate that hybrid feature selection enhances traditional machine learning classifiers, allowing them to achieve performance comparable to deep learning models while maintaining computational efficiency. This approach is particularly advantageous for deploying SER in resource-constrained environments, offering a balance between accuracy and efficiency in emotion recognition applications.

Read the paper · More papers on PaperTik