Determining an author's native language by mining a text for errors

Moshe Koppel, Jonathan Schler, Kfir Zigdon · 2005

In this paper, we show that stylistic text features can be exploited to determine an anonymous author's native language with high accuracy. Specifically, we first use automatic tools to ascertain frequencies of various stylistic idiosyncrasies in a text. These frequencies then serve as features for support vector machines that learn to classify texts according to author native language.

Read the paper · More papers on PaperTik