On Convergence of Discriminative Training Algorithm for Speaker Recognition
Srikanth R. Madikeri, Hema A. Murthy · 2011
In this paper, we present a discriminative training algorithm for GMM based speaker verification, and develop a convergence proof for the same. During training of the models, instead of performing MAP adaptation of the UBM, the model parameters of the GMM are estimated such that target scores are maximized while impostor scores are minimized. The focus of the algorithm is to estimate more accurately, the parameters of the elements of the mixture that are unique to a speaker. The algorithm uses an Expectation-Maximisation-like framework for estimation of parameters. It is shown that the algorithm converges with appropriate choice of a regularization parameter (α). α balances target and impostor data during each iteration of the algorithm. Fast convergence can be obtained by modifying the regularization parameter after each iteration. It is shown that incrementing α is a sufficient condition for convergence. Further, an impostor data selection technique for fast implementation is addressed. The performance of the system is evaluated on NIST 2003 benchmark dataset. A relative improvement 9.5% is obtained using Mel Frequency Slope as feature.