Author Identification Using Different Sizes Of Documents: A Summary
Samira Bourib, Salah Khennouf · Zenodo (CERN European Organization for Nuclear Research) · 2015
In the present research work, we deal with the problem of authorship attribution of ancient Arabic text documents, which were written by several ancient philosophers. For that purpose, we conducted several authorship attribution experiments applied with different text sizes. A special dataset, called “A4P” (Authorship Attribution for Ancient Arabic Philosophers), has been constructed by extracting texts of different sizes from the books of those 5 ancient Arabic philosophers, where the genre and topic are quite similar. The size of the texts varies from 100 words to 3000 words per text. In our approach two types of features are employed; character N-grams and words and several classifiers are used, namely: SMO based SVM, Multi Layer Perceptron, Linear Regression, Stamatatos distance and Manhattan distance. Results show that the minimum required text size (for getting good authorship attribution performances) depends on the used features and classification technique, but in the overall the performances of the proposed techniques are quite interesting.