Unsupervised pronunciation assessment analysis using utterance level alignment distance with self-supervised representations

Nayan Anand, Meenakshi Sirigiraju, Chiranjeevi Yarra · 2023

The pronunciation quality of second language (L2) learners can be affected by different factors including the following seven factors: Intelligibility, Intonation, Phoneme quality, mispronunciation, Mother tongue influence, Correct placement of pause, and Correct stress placement. An automatic assessment of these seven factors could be helpful for developing computer-assisted language learning systems. In this work, we assess the quality of all seven factors considering an unsupervised approach using DTW based utterance level alignment distance between expert’s and learner’s speech. Unlike the existing works that consider factor specific heuristic based features, we explore Wav2Vec-2.0 based self-supervised representations as a feature for assessing all the seven factors. The distance is computed using the following three distance metrics: Mean absolute error (MAE), Mean squared error (MSE), and Cosine distance (CD). Experiments are conducted on voisTUTOR corpus containing spoken English speech samples from 16 Indian L2 learners annotated with binary ratings (1 and 0) for all the factors. Using each distance metric, for each learner’s speech, four distance values are computed with a set of four expert samples. In the assessment, to circumvent the need for parallel expert data, we consider two (out of four) expert samples synthesized from the state-of-the-art text-to-speech (TTS) systems. We observe that the performance with the considered distance metric based unsupervised assessment approach is significantly more than that with the baseline for six out of seven factors under all three distance metrics and all four experts samples.

Read the paper · More papers on PaperTik