An N-gram Based Distributional Test for Authorship Identification

Kostas Fragos, Christos Skourlas · 2006

In this paper, a novel method for the authorship identification problem is presented. Based on character level text segmentation we study the disputed text's N-grams distributions within the authors' text collections. The distribution that behaves most abnormally is identified using the Kolmogorov -Smirnov test and the corresponding Author is selected as the correct one. Our method is evaluated using the test sets of the 2004 ALLC/ACH Ad-hoc Authorship Attribution Competition and its performance is comparable with the best performances of the participants in the competition. The main advantage of our method is that it is a simple, not parametric way for authorship attribution without the necessity of building authors' profiles from training data. Moreover, the method is language independent and does not require segmentation for languages such as Chinese or Thai. There is also no need for any text preprocessing or higher level processing, avoiding thus the use of taggers, parsers, feature selection strategies, or the use of other language dependent NLP tools.

Read the paper · More papers on PaperTik