Role of Morpho-Syntactic Features in Estonian Proficiency Classification
Sowmya Vajjala, Kaidi Lõo · 2013
• We developed an approach to perform proficiency classification for learners of Estonian as a second language. • Using a publicly accessible Estonian learner corpus, we show that – morpho-syntactic features in learner texts are useful predictors. – cascades of binary classifiers perform better than performing the classification in a single step. Related Work • SLA researchers studied the characteristic features of learner texts at different proficiency levels. (e.g., Tono, 2000; Vyatkina, 2012; Lu, 2012) • Automated assessment of student essays is also an active research area. (e.g., Yannakoudakis, Briscoe & Medlock, 2011; Burstein, 2013) • Contemporary research primarily focused on learner errors across proficiency levels. (e.g., Dickinson, Kübler & Meyer, 2012) • But, the role of morpho-syntactic features in proficiency classification was not explored before. Estonian Morphology • Estonian is agglutinative. Word forms can be formed by joining the morphemes together.- e.g., jalgades –>jalga+de+s (stem for foot +plural marker+inessive case marker) • It is fusional i.e., word forms can be formed by changing the stem.- e.g., jalg (foot, nominative), jala (genitive), jalga (partitive) • It has 14 productive cases (grammatical and semantic cases).- Cases express relations between words and are sometimes used instead of postpositions (jalal and jala peal have the same meaning: on the foot) • Cases have different alternative case endings.- e.g., Valid allative plural forms for jalg (foot) are: jalgadele, jalule, jalgele- We model some of these morphological characteristics as features for the learner proficiency classification task. The Corpus • The Estonian Interlanguage Corpus (EIC) consists of texts written by learners of Estonian as a Second Language (Eslon, 2007). • It mainly consists of short answers, essays and personal letters. • It also has error annotations but we did not use them in this paper. • Here is a numeric description of the corpus: