Realization and improvement algorithm of GMM - UBM model in voiceprint recognition

Jing Zhang · 2018

For the speaker recognition algorithm, the necessary training voice for the traditional voiceprint recognition model based on Gaussian mixture model (GMM) usually takes tens of seconds, and it is difficult to cover all the language features of the speaker, thus causing the phenomenon of low recognition rate. The paper proposed Gaussian mixture model- Universal background model (GMM-UBM), by which to train the common features and the proprietary features respectively. A large number of speaker voice was taken as a training sample, and the voice features covered most of the voice features was trained into a Universal background model, then the voice feature of specific human could be adaptively obtained through the Universal background model, so that a model with good performance could be trained out from a shorter voice, and the length of the voice was reduced and the recognition accuracy was improved. At the same time, in order to improve the accuracy, the improvement and solution to select the initial value and determine the mixing degree of GMM-UBM model were proposed. After the improvement, the test speech length was controlled at 4-7s, and the recognition rate could still reach 96%.

Read the paper · More papers on PaperTik