Non-parametric vector quantization of excitation source information for speaker recognition
Debadatta Pati, S. R. Mahadeva Prasanna · 2008
The objective of this work is to demonstrate the feasibility of excitation source information obtained by non-parametric vector quantization (VQ) for speaker recognition task. Linear prediction (LP) residual is used as the representation of excitation source information. The LP residual is subjected to non-parametric VQ during training. The codebooks are built for different codebook sizes. The testing of these codebooks using the LP residual of testing speech data indeed demonstrates that a codebook of sufficiently large size uniquely represents the speaker and provides appreciable performance. The speaker recognition system built using conventional Mel frequency cepstral coefficients (MFCCs) representing vocal tract information combines well with the proposed speaker recognition system using excitation source information to provide improved performance. On a set of randomly chosen 30 speakers from the TIMIT database, the proposed system provides 75%, MFCC based system provides 95% and the combined one provides 98.33%.