Speaker recognition in the case of emotional environment using transformation of speech features

Shashidhar G. Koolagudi, Shan E. Fatima, K. Sreenivasa Rao · 2012

This paper attempts speaker recognition from emotional speech by transforming speech features. Spectral features (MFCCs) are used to represent speaker specific emotional information. Simulated emotional Telugu database is used as speech data corpus to model the speakers. Gaussian mixture models (GMMs) are used to develop speaker recognition models. Eight different emotions namely anger, disgust, fear, happiness, neutral, sadness, sarcastic and surprise are used for characterizing speaker specific information. The average speaker recognition performance using neutral speech for both training and testing is observed to be 99.67%. This recognition performance reduces considerably when emotional sentences are used for testing. Results obtained in this work highly indicate the influence of emotional speech during speaker recognition.

Read the paper · More papers on PaperTik