A Multi-model SVR Approach to Estimating the CEFR Proficiency Level of Grammar Item Features
Brendan Flanagan, Sachio Hirokawa, Emiko Kaneko, Emi Izumi, Hiroaki Ogata · 2017
Analysis of publicly available language learning corpora can be useful for extracting characteristic features of learners from different proficiency levels. This can then be used to support language learning research and the creation of educational resources. In this paper, we classify the words and parts of speech of transcripts from different speaking proficiency levels found in the NICT-JLE corpus. The characteristic features of learners who have the equivalent spoken proficiency of CEFR levels A1 through to B2 were extracted by analyzing the data with the support vector machine method. In particular, we apply feature selection to find a set of characteristic features that achieve optimal classification performance, which can be used to predict spoken learner proficiency.