Can characters reveal your native language? A language-independent approach to native language identification
Radu Tudor Ionescu, Marius Popescu, Aoife Cahill · 2014
A common approach in text mining tasks such as text categorization, authorship identification or plagiarism detection is to rely on features like words, part-of-speech tags, stems, or some other high-level lin-guistic features. In this work, an approach that uses character n-grams as features is proposed for the task of native language identification. Instead of doing standard feature selection, the proposed approach combines several string kernels using mul-tiple kernel learning. Kernel Ridge Re-gression and Kernel Discriminant Analy-sis are independently used in the learning stage. The empirical results obtained in all the experiments conducted in this work in-dicate that the proposed approach achieves state of the art performance in native lan-guage identification, reaching an accuracy that is 1.7 % above the top scoring system of the 2013 NLI Shared Task. Further-more, the proposed approach has an im-portant advantage in that it is language in-dependent and linguistic theory neutral. In the cross-corpus experiment, the proposed approach shows that it can also be topic independent, improving the state of the art system by 32.3%. 1