Speaker verification using sparse representations on total variability i-vectors

Ming Li, Xiang Zhang, Yonghong Yan, Shrikanth Shri Narayanan · 2011

In this paper, the sparse representation computed by l1-minimization with quadratic constraints is employed to model the i-vectors in the low dimensional total variability space af-ter performing the Within-Class Covariance Normalization and Linear Discriminate Analysis channel compensation. First, we propose the background normalized l2 residual as a scoring cri-terion. Second, we demonstrate that the Tnorm can be effi-ciently achieved by using the Tnorm data as the non-target sam-ples in the over-complete dictionary. Finally, by fusing with the conventional i-vector based support vector machine (SVM) and cosine distance scoring system, we demonstrate overall system performance improvement. Experimental results show that the proposed fusion system achieved 4.05 % (male) and 5.25 % (fe-male) equal error rate (EER) after Tnorm on the single-single multi-language handheld telephone task of NIST SRE 2008 and outperformed the SVM baseline by yielding 7.1 % and 4.9 % rel-ative EER reduction for the male and female tasks, respectively. Index Terms: speaker verification, sparse representation i-vector modeling

Read the paper · More papers on PaperTik