Analysis of source and system features for speaker recognition in emotional conditions

K.N.R.K. Raju Alluri, Vishnu Vidyadhara Raju, Suryakanth V. Gangashetty, Anil Kumar Vuppala · 2016

Source and system features are extensively used in building Speaker Recognition (SR) systems. In this paper, we investigate the influence of source and system features on the performance of the SR system in emotional conditions. The Linear Prediction Residual Cepstral Coefficients (LPRCC) which corresponds to source features, and Mel Frequency Cepstral Coefficients (MFCC) and Linear Prediction Cepstral Coefficients (LPCC) that correspond to system features are used for modeling SR system. A maximum-likelihood classifier based on Gaussian mixture density functions is used and experiments are carried out on 3 standard emotional speech databases (Indian Institute of Technology-Simulated Emotion speech corpus (IITKGP-SESC): Hindi, IITKGP-SESC: Telugu and German Emotional Speech Database (EMO-DB)). The speaker models are trained with neutral utterances and tested with 3 types (anger, happy and sad) of emotional speech utterances. Based on experimental results, the performance degradation was observed in mismatched (trained with one emotion and tested with other emotion) case and the average percentage of degradation in SR task using source features is approximately 16% more compared to system features. The performance degradation of SR system is nullified when trained and tested with the same emotion.

Read the paper · More papers on PaperTik