Frequent Collocations and Authorial Style
David L. Hoover · Literary and Linguistic Computing · 2003
This paper examines the effectiveness of multivariate analysis of the frequencies of frequent collocations in characterizing authorial style. Cluster analyses of collocations over various spans, types, and linkages are performed on groups of texts by known authors to determine how well the frequencies of those collocations correctly attribute the texts to their authors and distinguish them from texts by other authors. In each case the results are compared with those based on the frequencies of frequent words and the frequencies of frequent sequences of words. Cluster analyses based on frequent words and sequences ascribe many of the texts to their correct authors. However, analyses based on frequent collocations are more accurate for several groups of texts, sometimes producing more completely correct attributions than analyses based on either words or sequences and sometimes producing the only completely correct attributions. They also produce results for small groups of problematic novels and critical texts extracted from the larger corpora that are often superior to those based on frequent words or frequent sequences. Finally, they perform better than analyses based on frequent words or sequences in simulated authorship attribution scenarios. Cluster analysis based on frequent collocations provides a robust and effective method of authorship attribution.