Text-independent speaker recognition by combining speaker-specific GMM with speaker adapted syllable-based HMM

Seiichi Nakagawa, W. Zhang, Mitsuo Takahashi · 2004

We presented a new text-independent speaker recognition method by combining a speaker-specific Gaussian mixture model (GMM) with a syllable-based HMM adapted by MLLR or MAP (S. Nakagawa et al., Proc. Eurospeech, p.3017-3020, 2003). The robustness of this speaker recognition method for speaking style changes was evaluated in this paper. A speaker identification experiment, using an NTT database, which consists of sentences of data uttered at three speed modes (normal, fast and slow) by 35 Japanese speakers (22 males and 13 females) on five sessions over ten months, was conducted. Each speaker uttered only 5 training utterances (about 20 seconds in total). We obtained an accuracy of 98.8% for text-independent speaker identification for three speaking style modes (normal, fast, slow) by using a short test utterance (about 4 seconds). This result was superior to conventional methods for the same database. We show that the attractive result was brought from the compensational effect between speaker specific GMM and speaker adapted syllable based HMM.

Read the paper · More papers on PaperTik