Local n-grams for author identification: notebook for PAN at CLEF 2013
Robert Layton, Paul Watters, Richard Dazeley · Cross-Language Evaluation Forum · 2013
Our approach to the author identification task uses existing authorship attribution methods using local n-grams (LNG) and performs a weighted ensemble. This approach came in third for this year’s competition, using a relatively simple scheme of weights by training set accuracy. LNG models create profiles, consisting of a list of character n-grams that best represent a particular author’s writing. The use of a weighted ensemble improved upon the accuracy of the method without reducing the speed of the algorithm; the submitted solution was not only near the top of the leaderboard in terms of accuracy, but it was also one of the faster algorithms submitted. The authorship identification task at PAN 2013 was a variation on a standard authorship analysis task of authorship attribution. In authorship identification, we have a training set of documents from the same author and a test document of unknown authorship. The task is to determine whether the author of the training documents was the one that wrote the test document. This task is different from authorship attribution in a few ways. First, we cannot simply take a ‘best guess’ whereby we find the best matching author. A decision on match or no match must be made, similar to the open set problem of authorship attribution, whereby the actual author may not be in the candidate set. Second, we have no point of reference to compare the similarity of author to document. In other words, we cannot know relatively if two profiles are similar and must therefore find algorithms that are able to know absolutely if two profiles match. Third, specifically for this task, the number of documents was small. Most problems in this task had just three documents from the same author, reducing the ability to determine variance.