Speaker recognition based on dynamic time warping and Gaussian mixture model

Nannan Zhang, Yanru Yao · 2020

At present, most of the speaker recognition models are based on the MFCC cepstrum feature of the mixture Gaussian model, because MFCC represents the speaker's voice characteristics, but the voice between different speakers is easy to be imitated. In view of this, this paper adds the pitch period to this model, and uses the combination of the two for speaker recognition. The pitch period contains the voice frequency structure information. Although it will be affected by the speaker's health, it is not easy to be imitated. At the same time, in view of the slow recognition speed caused by the conventional direct de mixing of Gaussian mixture model, this paper first uses DTW to calculate the shortest distance of pitch period between speech samples, and then uses GMM to calculate the maximum likelihood probability of distribution of test samples in the first few training samples with small scores. Through the experimental analysis, it can be seen that the speaker recognition model combined with DTW and GMM can improve the recognition speed. The accuracy and the reduction of recognition time are significant.

Read the paper · More papers on PaperTik