Study of MFCC and IHC Feature Extraction Methods With Probabilistic Acoustic Models for Speaker Biometric Applications

Aysha Sithara, Abraham Thomas, Dominic Mathew · Procedia Computer Science · 2018

Voice is an important human trait in natural human-to-human interaction / communication for identifying a person. So voice can be regarded as a biometric measure for recognizing or identifying the person similar to other biometric measures such as face, iris and fingerprints. Speaker recognition is a class of voice recognition where speaker is identified from the speech rather than the message. Automatic speaker recognition (SR) is an approach to identify people based on features extracted from speech utterances. The major task in any speaker recognition is to extract useful features and allow meaningful patterns of speaker models. This paper compares the performance of two feature extraction techniques Mel Frequency Cepstral Coefficient (MFCC) and Inner Hair Cell Coefficient (IHC) with two different modelling methods Gaussian Mixture Model - Universal background model (GMM - UBM) and i- vector approach. In this experiment speech samples of 600 speakers from TIMIT database with 10 utterances of each speaker are taken for identifying the speaker. A text independent speaker recognition system was implemented and this study resulted in an inference of MFCC feature outperforms IHC feature for both GMM and i vector. The performance of voiced speech IHC feature which simulates the physiological behaviour of human ear is better in terms of accuracy than full speech (voiced and unvoiced) in GMM.

Read the paper · More papers on PaperTik