Native Language Identification Using a Mixture of Character and Word N-grams
Elham Mohammadi, Hadi Veisi, Hessam Amini · 2017
Native language identification (NLI) is the task of determining an author's native language, based on a piece of his/her writing in a second language.In recent years, NLI has received much attention due to its challenging nature and its applications in language pedagogy and forensic linguistics.We participated in the NLI Shared Task 2017 under the name UT-DSP.In our effort to implement a method for native language identification, we made use of a mixture of character and word Ngrams, and achieved an optimal F1-score of 0.7748, using both essay and speech transcription datasets.