DA-IICT Cross-lingual and Multilingual Corpora for Speaker Recognition

Hemant A. Patil, Sunayana Sitaram, Esha Sharma · 2009

In this paper the design and development of the DA-IICT Cross-lingual and Multilingual Speech Corpora is presented which includes unconventional sounds like cough, whistle, whisper, frication, idiosyncrasies, etc. from bilingual subjects (i.e., who can speak Hindi and Indian English) and trilingual subjects (who can speak Hindi, Indian English and mother tongue) for the development of Automatic Speaker Recognition System. Thirteen Indian languages and the Nepali language are considered as the subjects’ mother tongue/native languages. Unconventional sounds are considered to examine how much speaker-specific information they carry. Finally, an ASR system based on spectral or cepstral features (i.e., LPC, LPCC, MFCC) and polynomial classifier of 2nd order approximation is presented to evaluate the developed corpora.

Read the paper · More papers on PaperTik