Mind your corpus: systematic errors in authorship attribution

M. Eder · Literary and Linguistic Computing · 2013

In computational stylistics, any influence of unwanted noise—e.g. caused by an untidily prepared corpus—might lead to biased or false results. Relying on contaminated data is similar to using dirty test tubes in a laboratory: it inescapably means falling into systematic error. An important question is what degree of nonchalance is acceptable to obtain sufficiently reliable results. The present study attempts to verify the impact of unwanted noise in a series of experiments conducted on several corpora of English, German, Polish, Ancient Greek, and Latin prose texts. In 100 iterations, a given corpus was gradually damaged, and controlled tests for authorship were applied. The first experiment was designed to show the correlation between a dirty corpus and attribution accuracy. The second was aimed to test how disorder in word frequencies—produced by scribal and/or editorial modifications—affects the attribution abilities of particular corpora. The goal of the third experiment was to test how much ‘authorial’ data a given text needs to have to trace authorial fingerprint through a mass of external quotations.

Read the paper · More papers on PaperTik