Frequency warping and robust speaker verification: a comparison of alternative mel-scale representations
Tomi Kinnunen, Md. Jahangir Alam, Pavel Matějka, Patrick J Kenny, Jaň Černocký, Douglas D. O’Shaughnessy · 2013
Accuracy of speaker verification is high under controlled condi-tions but falls off rapidly in the presence of interfering sounds. This is because spectral features, such as Mel-frequency cep-stral coefficients (MFCCs), are sensitive to additive noise. MFCCs are a particular realization of warped-frequency rep-resentation with low-frequency focus. But there are several alternative, potentially more robust, warped-frequency repre-sentations. We provide an experimental comparison of five warped-frequency features. They use exactly the same fre-quency warping function, the same number of coefficients and postprocessing, but differ in their internal computations. The compared variants are (1) conventional MFCCs from discrete Fourier transform (DFT), followed by Mel-scaled filterbank, (2) MFCCs via direct warping of DFT, followed by linear-scale fil-terbank, (3) warped linear prediction features, (4) perceptual minimum variance distortionless features and (5) recently pro-posed sparse Mel-scale histogram features. Experiments car-ried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases. Index Terms: speaker recognition, noise, frequency warping 1.