Researches based on neural network about optimizing combination of speaker verification

Tiantian Chenm, Gang Liu · 2017

The identity vector (i-vector) approach has been the state-of-the-art for text-independent speaker recognition, both identification and verification in recent years. An I-vector is a low-dimensional vector, which is called total variability space and is represented with a thin and tall rectangular matrix. In this paper, there is a novel interpretation of the Universal Background Model (UBM), and consider it as a mapping function that transforms the variable length observations (speech utterances) into a fixed dimensional feature vector (sufficient statistics). After this mapping, a similarity measurement is computed on the fixed dimensional features. With this novel interpretation, we proposed new combinations which produce improvements over the conventional UBM framework in both equal error rate and detection cost function. Performance can be further improved by progressively combining different kinds of the feature and the network construction in this vector space via the increasing numbers of neural network layers. Our algorithm makes use of a hybrid generative-discriminative framework: it uses a generative model to learn the characteristics of a speaker and then a discriminative model to discriminate between a speaker and an impostor.

Read the paper · More papers on PaperTik