Voistutor 2.0: A Speech Corpus with Phonetic Transcription for Pronunciation Evaluation of Indian L2 English Learners
Priyanshi Pal, Chiranjeevi Yarra, Prasanta Ghosh · 2022
In computer assisted pronunciation training (CAPT), robust automatic models are critical for pronunciation assessment and mispronunciation detection and diagnosis (MDD). In the modelling, besides the audio data of second language (L2) learners, CAPT requires manually annotated ratings of overall pronunciation quality, and the MDD uses manually annotated phonetic transcriptions. Though the pronunciation quality and the mispronunciation are interdependent, to the best of our knowledge, none of the existing corpora contains both ratings and phonetic transcriptions. This could be due to the cost involved in obtaining phonetic transcriptions. However, a corpus with both kinds of information could benefit the researchers to obtain robust models by exploring the interdependencies. For addressing this, we develop voisTUTOR 2.0 corpus considering the existing voisTUTOR corpus referred to as voisTUTOR 1.0. We obtain phonetic transcriptions manually from a linguist for the entire Indian L2 learners’ English audio data (26529 utterances approximately 14 hours) in voisTUTOR 1.0 for which overall quality ratings and binary scores of factors influencing the pronunciation quality are available. A preliminary analysis of voisTUTOR 2.0 suggests that the phonetic errors correlated with the ratings and the binary scores indicating mispronunciations and phoneme quality.