An Automated System for Regional Nativity Identification of Indian speakers from English Speech
Radha Krishna Guntur, R. Krishnan, Vinay Kumar Mittal · 2019
This paper proposes an automated system to identify speaker's regional nativity by analysing their English speech utterances. A database of English speech of native speakers of three South Indian languages: Kannada (KAN), Tamil (TAM) and Telugu (TEL), is especially collected for this study, in text-independent mode. Mel Frequency Cepstral Coefficients (MFCCs) features are used with three different classifiers, namely, Gaussian Mixture Model (GMM), GMM-Universal Background Model (GMM-UBM) and i-vector. The i-vector classifier gave accuracies of 93.9%. Nativity identification from English speech is observed to be relatively easier for native speakers of Kannada language, than for Tamil and Telugu speakers.