Sentence selection for automatic scoring of Mandarin proficiency

Jiahong Yuan, Xiaoying Xu, Wei Lai, Weiping Ye, Zhao Xinru, Mark Yoffe Liberman · 2015

A central problem in research on automatic proficiency scoring is to differentiate the variability between and within groups of standard and non-standard speakers. Along with the effort to improve the robustness of techniques and models, we can also select test sentences that are more reliable for measuring the between-group variability. This study demonstrated that the performance of an automatic scoring system could be significantly improved by excluding “bad” sentences from the scoring procedure. The experiments on a dataset of Putonghua Shuiping Ceshi (Mandarin proficiency test) showed that, compared to all available sentences, using only best-performed sentences improved the speaker-level correlation between human and automatic scores from r = .640 to r = .824.

Read the paper · More papers on PaperTik