Initial Analysis of the Impact of Emotional Speech on the Performance of Speaker Recognition on New Serbian Emotional Database
Igor Mandaric, Mia Vujović, Siniša Suzić, Tijana Nosek, Nikola Simić, Vlado Delić · 2021 29th Telecommunications Forum (TELFOR) · 2021
In the paper we compared three ML methods (kNN, SVM and MLP) to build the optimal models for speaker recognition for two datasets with different recording conditions. We studied the impact of different speech features on classification performance, with the main focus given to MFCCs. All models were built using neutral speech, but their performance on emotional test data is also analyzed. The achieved accuracy on speech in neutral style was ~99% for SEAC dataset and ~97% for VCTK. We observed a significant decrease in the results on emotional data. An improvement occured when other features from Interspeech 2009 feature set were added to MFCC in the model creation.