A novel dynamic acoustical model for speaker verification

Gongjun Li, Carol Espy-Wilson · The Journal of the Acoustical Society of America · 2004

In speaker verification, the conventional acoustical models (hidden Markov model and vector quantization) are not able to capture a speaker’s dynamic characteristics. In this paper we describe a novel dynamic acoustical model. The training data are viewed as a concatenation of many speech-pattern samples, and the pattern matching involves a comparison of the pattern samples and the test speech. To reduce the amount of computation, a tree is generated to index the entrance to pattern samples using an expectation and maximization (EM) approach, and leaves in the tree are employed to quantize the feature vectors in the training data. The obtained leaf-number sequences are exploited in pattern matching as a temporal model. We use a DTW scheme and a GMM scheme to match the training data and the test speech. Experimental results on NIST’98 speaker recognition evaluation data show that the accuracy of speaker verification on 3- and 10-s test speech is raised from 71.1% and 75.2% for a baseline GMM-based system to 80.0% and 82.1% for the dynamic acoustical model, respectively. Furthermore, some pattern samples in the training data are correctly tracked by the test speech.

Read the paper · More papers on PaperTik