Significance of utterance partitioning in GMM-SVM based speaker verification in varying background environment
Sourjya Sarkar, K. Sreenivasa Rao · 2013
This paper explores the GMM-SVM combined approach for text-independent speaker verification in rapidly varying environmental noise. For mitigating the effect of mismatched training and test utterance length and countering the impact of data-imbalance in SVM scoring, we partition each full-length enrollment utterance into a number of sub-utterances and derive a GMM supervector from each of them prior to SVM training. Experiments conducted on the NIST-SRE-2003 database demonstrate that the GMM-SVM system with partitioned utterances outperform the conventional GMM-SVM based speaker verification system. The noisy background is simulated by degrading non-overlapping segments of each training and test utterance of the NIST-SRE-2003 by additive noises (car, factory, pink & white) collected from the NOISEX-92 database, at OdB, 5dB, 7dB & 10dB SNRs respectively. An average performance improvement of 6.41% and 10.56% EER across all SNRs is observed in comparison to the traditional GMM-SVM and GMM-UBM based systems, respectively.