Intonational speaker verification: A study on parameters and performance under noisy conditions

Sadjad Siddiq, Tomi Kinnunen, Martti Vainio, Stefan Werner · 2012

Prosody-based speaker verification using fundamental frequency (f0) is considered. Our study consists of two phases. First, we do extensive optimization of parameters to establish a baseline system before dealing with noisy conditions. This includes a study of f0extractor parameters, choice of features (discrete cosine transform, discrete Fourier transform, Legendre polynomials, linear prediction), f0track interpolation (none, linear, Hermite), framing parameters and windowing (none, Hamming), f0representation domain (linear, log), number of transformation coefficients and, finally, use of higher-level delta coefficients. Using the optimized parameters, we then explore the robustness of prosody features under white noise and factory noise degradations. Using a GMM-UBM system on the NIST 2006 SRE corpus, we reach an EER of 28.4 % and 27.6 % for the intonational and MFCC features respectively at -20 dB SNR white noise contamination; fusion of the two yields an EER of 24.38 %.

Read the paper · More papers on PaperTik