Assigning CEFR-J levels to English texts based on textual features
Satoru Uchida, Masashi Negishi · Kyushu University Institutional Repository (QIR) (Kyushu University) · 2018
The present study attempts to assign CEFR-J levels (Pre-A1 to C2) to English texts based on textural features. Based on a coursebook corpus which consists of EFL/ESL English textbooks that claim to be based on CEFR, four textual indexes are calculated. The indexes are ARI (a readability measure), VperSent (an average number of verbs included in each sentence), AvrDiff (the average of word difficulties) and BperA (the ratio of B level content words to A level content words). Regression models are created for each index to predict the level of the input text which is then implemented as an online application called CVLA (CEFR-based Vocabulary Level Analyzer). To show how CVLA works, experiments are conducted using major English language ability tests.