Auditory model representation for speaker recognition

John M. Colombi, T.R. Anderson, Steven K. Rogers, D.W. Ruck, Gregory T. Warhola · IEEE International Conference on Acoustics Speech and Signal Processing · 1993

An examination of the KING database that compares proven spectral processing techniques with an auditory model representation for speaker recognition is presented. The feature sets compared are LPC (linear predictive coding) cepstral coefficients and auditory nerve firing rates provided by the Payton model. The two feature sets were quantized by two clustering algorithms, a Linde-Buzo-Gray algorithm and a Kohonen self-organizing feature map. The resulting vector quantized distortion based classification indicates that the auditory model provides accuracies comparable with LPC cepstral in nonstudio quality environments and over multiple sessions. For a 10-speaker subset using only voiced frames of 15-s segments, both achieve over 80% identification rate. Cepstral performs better on verification tasks measured with receiver operating characteristics curves.>

Read the paper · More papers on PaperTik