Advance statistical modeling for speaker verification
Rupali Waghmare, Vinayak G. Asutkar · 2008
Speaker verification deals with the problem of verifying whether a given utterance has been pronounced by a claimed authorized speaker. This problem is important because an accurate speaker verification system can be applied to many security applications. Fundamental frequency is the rate of vocal folds vibration during speech. It is considered to be one of the most important prosodic features to characterize speech and speaker specific patterns. This study considers two aspects: the estimation of fundamental frequency and its modeling for speaker recognition. Prosodic information has been applied in two main ways are global statistics and dynamic time warping. To understand this prosodic parameter, we present an analysis of three popular algorithms: correlation, kurtosis and kernel function. In this paper, we present a new model for speaker verification without background model, which is called correlation and kernel function method (CK method). In CK method, the correlation and un-correlation of MFCC are used to identify individuals, and a kernel function is used to work out the likelihood of two models. This method is faster than GMM method, requires fewer data to train and also less space to store the model. From experimental results we say that the kernel function has higher accuracy than correlation and kurtosis method.