Speaker recognition in duration-mismatched condition using bootstrapped i-vectors

Atsushi Ando, Taichi Asami, Yoshikazu Yamaguchi, Yushi Aono · 2016

This paper presents a novel speaker recognition framework that handles duration mismatch between registered and test utterances. The i-vectors extracted from short utterances exhibit high variance due to phoneme imbalance, which causes performance degradation in the duration mismatch condition. Most conventional methods attempt to decrease the variance by offsetting i-vectors or speaker similarity scores, however, the variances caused by duration differences are usually too complex to offset. Instead of conventional offsetting approaches, our proposed method, inspired by ensemble learning, attains low-variance results by generating multiple fixed-length short utterances from registered/test utterances and integrating their speaker similarities. Bootstrapped i-vectors are yielded from generated short utterances and the average PLDA scores between the combinations of registered and test bootstrapped i-vectors are used for speaker decision. Experiments show that the proposed method improves the equal-error-rate in trials of 60 second registered utterances and less than 5 second test utterances with relative error reduction of 73.9%-90.6%. Moreover, it appears that the proposed method has smaller score variance than the baseline.

Read the paper · More papers on PaperTik