Gradual enhancement of GMM for speaker identification

Juraj Kačur · 2017

The article focuses on different initialization and training strategies for Gaussian Mixture Model (GMM) applied to speaker identification domain. The emphasis is on the overall performance of GMM in the text independent speaker identification task. Stochastic initialization scheme selecting centroids randomly from among training vectors was augmented by gradual mixture splitting using different strategies interlaced with training cycles. In the training phase several structures and training approaches like ML training and adaptation of means or whole models were tested and evaluated using standard acoustical Mel frequency features. The evaluation and training tests were accomplished on a speaker database containing 2 environments, each with training and testing parts. It was shown that the best results were achieved by using universal background model (UBM) of a speaker for adapting whole models of individuals as opposed to more common approach based on the mean MAP adaptation. In the case of nonexistence of UBM the gradual splitting strategies proved to be superior to random initialization approaches both in match and mismatch training-testing conditions.

Read the paper · More papers on PaperTik