Norwegian Native Language Identification
Shervin Malmasi, Mark Dras, Irina Temnikova · 2015
We present a study of Native Language Identification (NLI) using data from learn-ers of Norwegian, a language not yet used for this task. NLI is the task of predicting a writer’s first language using only their writings in a learned language. We find that three feature types, function words, part-of-speech n-grams and a hy-brid part-of-speech/function word mixture n-gram model are useful here. Our sys-tem achieves an accuracy of 79 % against a baseline of 13 % for predicting an author’s L1. The same features can distinguish non-native writing with 99 % accuracy. We also find that part-of-speech n-gram per-formance on this data deviates from previ-ous NLI results, possibly due to the use of manually post-corrected tags. 1