A Comparison of Various Adaptation Methods for Speaker Verification With Limited Enrollment Data

Man‐Wai Mak, Roger Wend-Huu Hsiao, Brian Kan-Wing Mak · 2006

One key factor that hinders the widespread deployment of speaker verification technologies is the requirement of long enrollment utterances to guarantee low error rate during verification. To gain user acceptance of speaker verification technologies, adaptation algorithms that can enroll speakers with short utterances are highly essential. To this end, this paper applies kernel eigenspace-based MLLR (KEMLLR) for speaker enrollment and compares its performance against three state-of-the-art model adaptation techniques: maximum a posteriori (MAP), maximum-likelihood linear regression (MLLR), and reference speaker weighting (RSW). The techniques were compared under the NIST2001 SRE framework, with enrollment data vary from 2 to 32 seconds. Experimental results show that KEMLLR is most effective for short enrollment utterances (between 2 to 4 seconds) and that MAP performs better when long utterances (32 seconds) are available.*This work was supported by the Research Grant Council of the Hong Kong SAR (Project Nos. CUHK 1/02C and PolyU 5214/04E).†Roger completed this work while he was with the Hong Kong University of Science and Technology before he left for CMU.‡This research is partially supported by the Research Grants Council of the Hong Kong SAR under the grant number CA02/03.EG04.

Read the paper · More papers on PaperTik