Native Language Identification: A Key N-gram Category Approach

Kristopher Kyle, Scott A. Crossley, Jianmin Dai, Danielle S. McNamara · 2013

This study explores the efficacy of an approach to native language identification that utilizes grammatical, rhetorical, semantic, syntactic, and cohesive function categories comprised of key n-grams. The study found that a model based on these categories of key n-grams was able to successfully predict the L1 of essays written in English by L2 learners from 11 different L1 backgrounds with an accuracy of 59%. Preliminary findings concerning instances of crosslinguistic influence are discussed, along with evidence of language similarities based on patterns of language misclassification. 1.

Read the paper · More papers on PaperTik