A simple stylometric comparator: nifty assignment

Steven Benzel · Journal of computing sciences in colleges · 2015

Stylometry is the study of linguistic style and a common concern is the determination of authorship of a written work. We present a very straightforward syntactical comparator for two works which is surprisingly useful in predicting authorship. The comparator is simple enough that it can be introduced in an introductory data structures class. To begin, let N be a small positive integer, τ the set of tokens consisting of the words in the English language, and T a stream of such tokens, for example a text in English stripped of punctuation. We can then construct an associative array A with key: value pairs consisting of token sequences of length N along with their frequency in T. For example, if we take N=3 with the text Huckleberry Finn by Mark Twain and we sort A by highest to lowest values, the first 10 entries of the array will be:

Read the paper · More papers on PaperTik