Beyond Native Norms: A Perceptually Grounded and Fair Framework for Automatic Speech Assessment

Mewlude Nijat, Yang Wei, Shengze Li, Abdusalam Dawut, Askar Hamdulla · Applied Sciences · 2026

Pronunciation assessment is central to computer-assisted pronunciation training (CAPT) and speaking tests, yet most systems still adopt a native norm, treating deviations from canonical L1 pronunciations as errors. In contrast, rating rubrics and psycholinguistic evidence emphasize intelligibility for a target listener population and show that listeners rapidly adapt their phonetic categories to new accents. We argue that automatic assessment should likewise be referenced to the target learner group. We build a Transformer-based mispronunciation detection (MD) model that computationally mimics listener adaptation: it is first pre-trained on multi-speaker Librispeech, then fine-tuned on the non-native L2-ARCTIC corpus that represents a specific learner population. Fine-tuning, using either synthetic or human MD labels, constrains updates to the phonetic space (i.e., the representation space used to encode phone-level distinctions, the learned phone/phonetic embedding space, and its alignment with acoustic representations), which means that only the phonetic module is updated while the rest of the model stays fixed. Relative to the pre-trained model, L2 adaptation substantially improves MD recall and F1, increasing ROC–AUC from 0.72 to 0.85. The results support a target-population norm and inform the design of perception-aligned, fairer automatic pronunciation assessment systems.

Read the paper · More papers on PaperTik