An iVector extractor using pre-trained neural networks for speaker verification

Shanshan Zhang, Rong Jian Zheng, Bo Xu · 2014

The iVector representation of speech utterances is currently widely used in speaker and language recognition tasks. In this paper, an iVector extractor using pre-trained neural networks is proposed for speaker verification. It can be viewed as an alternative to the classical total variability approach. In the proposed system, a neural network with bottleneck layer is trained with speaker labeled utterances, then we utilize the bottleneck features of the network to represent the input utterance. As a new iVector representation, it shows comparable performance with the state-of-the-art Total Variability Model (TVM) based iVector extraction system on NIST 2008 SRE. We further achieve a 10% reduction in equal error rates with combination of the proposed extraction system and the TVM system.

Read the paper · More papers on PaperTik