Enhancing Emotion Recognition in Voice: Leveraging Support Vector Machines
Rio Galang Jati Respati, Arrie Kurniawardhani, Irving Vitra Paputungan · 2024
Voice analysis has emerged as a pivotal tool in emotion detection, with applications spanning customer service to aviation safety. This study investigates the effectiveness of Support Vector Machine (SVM) classification for emotion detection from voice data, focusing on the impact of feature extraction methods and dataset characteristics. Utilizing Mel-Frequency Cepstral Coefficients (MFCCs) for feature extraction, we evaluated SVM performance with different kernel functions and regularization parameters across two datasets: the RAVDESS and JL Corpus. The RAVDESS dataset, featuring 2,452 recordings in English, includes eight emotional categories, while the JL Corpus, with 480 recordings, encompasses ten emotions. Our results reveal that an 80% training and 20% testing split provided the highest accuracy, precision, and F1-score. The Radial Basis Function (RBF) kernel with a regularization parameter$C$of 1000 outperformed the Linear kernel, achieving superior accuracy and faster training times. The JL Corpus dataset achieved an accuracy above 90%, while the RAVDESS dataset yielded 63% accuracy, primarily due to fluctuating amplitude patterns. Combining the datasets improved overall performance to nearly 70%, highlighting the benefits of diverse data in enhancing model generalization.