Combining MAP and MLLR Approaches for SVM Based Speaker Recognition with a Multi-class MLLR Technique
Haipeng Wang, Xiang Zhang, Xiang Xiao, Jianping Zhang, Yonghong Yan · 2009
Gaussian mixture models with an universal background model (UBM) have been the standard method for speaker recognition. Typically, maximum a posteriori (MAP) or maximum likelihood linear regression (MLLR) is used to adapt the means of the UBM. Together with the SVM modeling technique, these approaches can achieve excellent performance. MLLR is quite efficient when the amount of adaptation data is limited, but has poor asymptotic properties as the amount of data increases. MAP estimation has nice asymptotic properties, but provides only a moderate improvement when the amount of adaptation data is small. In this paper, in order to take advantage of both approaches to improve the recognition performance, a new approach for speaker adaptation consisting of MAP adaptation followed by MLLR adaptation is presented. This work is enriched by a multi-class MLLR technique, which clusters the Gaussian components into regression classes and applies a different transform to each class. Experiments on the NIST 2006 SRE corpus show that the proposed approach improves on both MLLR and MAP adaptation systems.