Syllable category based short utterance speaker recognition

Nakhat Fatima, Thomas Fang Zheng · 2012

In Short Utterance Speaker Recognition (SUSR), the role of complete speech units like syllables in carrying speaker information needs further investigation. This paper presents a novel method of using syllable categories for SUSR. We define Syllable Categories (SCs) with the help of syllable structure of Chinese language. Syllables in speech are segmented into SCs, which are then used to develop Universal Background SC Model for each SC. Conventional GMM-UBM system is used for training and testing. The proposed categories give average EER of 17.79%, 19.35% and 21.65% for 3, 2 and 1 second of test utterance length respectively. Experimental results show that in text dependent SUSR, significant speaker-specific information is present at syllable level where prosodic idiosyncrasies can be utilized. This information can be used in SUSR by exploiting similarities in consonants and vowels of a syllable such that SCs can be used effectively.

Read the paper · More papers on PaperTik