Cognate and Misspelling Features for Natural Language Identification

Garrett Nicolai, Bradley D. Hauer, Mohammad Yahya Bani Salameh, Lei Yao, Grzegorz Kondrak · Workshop on Innovative Use of NLP for Building Educational Applications · 2013

We apply Support Vector Machines to differentiate between 11 native languages in the 2013 Native Language Identification Shared Task. We expand a set of common language identification features to include cognate interference and spelling mistakes. Our best results are obtained with a classifier which includes both the cognate and the misspelling features, as well as word unigrams, word bigrams, character bigrams, and syntax production rules.

Read the paper · More papers on PaperTik