Adaptation of data fusion-based speaker verification models
Kevin R. Farrell · 2003
This paper presents methods for adapting models in a data fusion-based speaker verification system. The models that are used in the data fusion system are the neural tree network (NTN), dynamic time warping (DTW), and hidden Markov model (HMM). The models provide information based on discriminant information, distortion. measurements, and probabilistic evaluation, respectively. The parameters of these models are updated during the adaptation process using verification data. This allows the models to track changes in the users voice over time and additionally allows the technology to supplement the typically limited data obtained at enrollment. The adaptation algorithms are evaluated for both cases where the data is known to come from the correct user and not known to come from the correct user. For the case where the adaptation data is not known to come from the correct user, threshold criteria is used for determining if the adaptation should occur or not. Experiments were performed on voice data collected within landline telephony, wireless telephony, and multimedia environments. The adaptation lead to a 20% relative reduction on the equal error rate.