Automatic prediction of vocabulary knowledge for learners of Chinese as a foreign language
John Lee, Chak Yan Yeung · 2018
Since extensive reading is beneficial for learning a foreign language, students are encouraged to seek additional reading materials from sources beyond their textbooks. The materials should be difficult enough to stretch the student's language proficiency, but not too difficult as to hinder comprehension. A complex word identification (CWI) system can identify texts that optimize these criteria by estimating the student's proficiency level. We present a personalized CWI model for Chinese as a foreign language. This model predicts whether the student knows a Chinese word or not, based on a small training set from the student. In empirical evaluation, a support vector machine (SVM) classifier with features based on graded vocabulary lists yielded the best performance, outperforming a label propagation approach that is state-of-the-art for personalized CWI for English.