Dependency of recognition rate on number of words for text-independent speaker recognition using vector quantization
Hidenori Shimizu, Tetsuo Funada · The Journal of the Acoustical Society of America · 2008
In this research, we discuss speaker recognition using the Kohonen feature map. The map is constructed for each speaker, and it is trained by using log-power and fourteenth-order mel-frequency cepstral coefficients (MFCC) and their temporal difference. The quantization distortion is computed between the input speech and a specific vector on the feature map of each speaker. We conduct speaker recognition experiment based on VQ distortion. Utterances of prefectural name in Japan are used as speech data. We examine particularly the dependency of recognition rate on number of words used for recognition. According to our experiments of speaker identification, this system correctly recognizes 98.9% by using a single word for 40 male speakers, while it attains 100% by using more than three words. Moreover, we confirmed superiority of using VQ over HMM under the same experimental conditions.