Discriminative linear transforms for speaker adaptation
L.F. Uebel, PC Woodland · 2001
Linear transform adaptation techniques such as Maximum Like-lihood Linear Regression (MLLR) are a popular and effective family of methods for speaker adaptation. MLLR estimates transform parameters for Gaussian means and variances using a maximum likelihood (ML) objective function. This paper dis-cusses the use of an alternative discriminative objective func-tion for linear transform estimation, which is an interpolation of the maximum mutual information (MMI) objective function and the ML criterion. This Discriminative Linear Transform (DLT) more directly reduces the word error rate of the adaptation data than MLLR and assuming good generalisation will also reduce test-set error rates. The implementation of DLT estimation is discussed and test-data recognition results compared to those from standard unconstrained MLLR (mean and variance adap-tation) using the 1994 WSJ/NAB spoke 3 non-native adaptation task. The results show that relative reductions in word error rate between 7 % and 19 % can be obtained by using DLTs. 1.