Study on the varying degree of speaker identity information reflected across the different MFCCs
Saharul Alom Barlaskar, Mohammad Azharuddin Laskar, Nirupam Shome, Rabul Hussain Laskar · 2016
This paper presents a study on the varying degree of speaker identity information that is embedded across the MFCCs feature. A DTW based template matching model has been used on a set of 100 speakers' data, each from the CSLU speaker recognition database and the NITs text-dependent speaker verification database. The 39 dimensional MFCC feature vector, conventionally used in the automatic speaker recognition system is divided on an experiment basis into five sets. This is based on the understanding that the lower order MFCCs and the higher order ones represent significantly different characteristics of the vocal tract. Appropriate feature level channel compensation techniques of cepstral mean substraction (CMS) and spectral mean deviation (SMD) are also used in the study so as to mitigate the channel variability and offer better comparative analysis of the different feature sets. It has been observed that the set of the first six lower order co-efficients yield the best result in terms of EER. This indicates that the lower order co-efficients have higher degree of speaker identity information.