Predicting accentedness and comprehensibility through ASR scores and acoustic features
Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik · Computer Speech & Language · 2025
• Compared two automatic speech recognition models (TDNN and Whisper) for predicting accentedness and comprehensibility. • Used a data-driven method to select the most relevant acoustic features for accentedness and comprehensibility. • Used a linear mixed-effects model to incorporate speaker and utterance differences. • Combined segmental and suprasegmental features to better understand accentedness and comprehensibility of non-native speech. Accentedness and comprehensibility scales are widely used in measuring the oral proficiency of second language (L2) learners, including learners of English as a Second Language (ESL). In this paper, we focus on gaining a better understanding of the concepts of accentedness and comprehensibility by developing and applying automatic measures to ESL utterances produced by Indonesian learners. We extracted features both on the segmental and the suprasegmental (fundamental frequency, loudness, energy et al.) levels to investigate which features are actually related to expert judgments on accentedness and comprehensibility. Automatic Speech Recognition (ASR) pronunciation scores based on the traditional Kaldi Time Delay Neural Network (TDNN) model and on the End-to-End Whisper model were applied, and data-driven methods were used by combining acoustic features extracted by the Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) and Praat. The experimental results showed that Whisper outperformed the Kaldi-TDNN model. The Whisper model gave the best results for predicting comprehensibility on the basis of phone distance, and the best results for predicting accentedness on the basis of grapheme distance. Combining segmental and suprasegmental features improved the results, yielding different feature rankings for comprehensibility and accentedness. In our final step of analysis, we included differences between utterances and learners as random effects in a mixed linear regression model. Exploiting these information sources yielded a substantial improvement in predicting both comprehensibility and accentedness.