Multi-Source Domain Adaptation and Fusion for Speaker Verification
Donghui Zhu, Ning Chen · IEEE/ACM Transactions on Audio Speech and Language Processing · 2022
Recently, domain adaptation model was studied for Speaker Verification (SV) task to solve the performance reduction resulted from various mismatch between the training dataset and the testing dataset. However, since most schemes apply adaptation from single source domain to the target domain, they may not be robust enough against various mismatch between the samples in the source domain and those in the target domain. In this paper, multiple source domain adaptation and fusion model was studied to solve this problem. First, the x-vector based SV scheme is pretrained on each source domain. Second, UnShared Network (USN) combining weight regularization loss based adaptation model is adopted to adapt each pretrained model to the target domain and maintain the speaker-related feature as much as possible. Third, the Registered-Utterance-To-Test-Utterance (RUTTU) similarity matrix is constructed based on the x-vector feature extracted by the adapted model from each source domain. Fourth, the Similarity Network Fusion (SNF) technique is introduced and modified to obtain Modified Similarity Network Fusion (MSNF) scheme, which is used to fuse the RUTTU similarity matrices obtained from multiple adapted models. Finally, speaker matching is achieved by averaging the obtained fused similarities between the test utterance and each of the registered utterances of a specific speaker. Extensive experimental results demonstrate that i) the proposed fusion model outperforms the Multi-source Distilling Domain Adaptation (MDDA) model and Logistic Regression Fusion (LRF) model in SV task; ii) the USN combining weight regularization loss based adaptation strategy, the MSNF fusion scheme, and the modified speaker matching mechanism all contribute to the performance enhancement; iii) the fusion effectiveness is little influenced by the hyper-parameters setting of MSNF.