Combination of pitch and MFCC GMM supervectors for speaker verification

Wei Huang, Jianshu Chao, Yaxin Zhang · 2008

A large majority of speaker verification systems are based on frame-level acoustic features, such as Mel frequency cepstral coefficients (MFCCs) which characterize the vocal tract contribution. The most commonly used statistical GMM-UBM classifier models the distribution of MFCCs quite well. Pitch is one of the most important features which characterize speaker-dependent vocal fold vibration rate. It can complement the vocal tract information as source information. Although the source information is supposed to follow a lognormal distribution, the discriminative support vector machine (SVM) is more suitable for pitch classification. In this paper, firstly we exploit GMM-UBM and SVM to the frame-level pitch vectors. Then we put the state-of-the-art GMM supervectors concept to the pitch feature vectors and experiment shows a promising result. And the combination of two feature type GMM supervectors systems gains much better performance. All experiment results are obtained on the NIST 2001 speaker database.

Read the paper · More papers on PaperTik