Low Frequency Words, Genre, Date, and Authorship

T. Merriam · Notes and Queries · 2006

IN ‘Testing Burrow's Data’, David L. Hoover draws attention to evidence that the power of authorship discrimination increases with the inclusion of large numbers of low frequency or ‘content words’.1 For years analysts have eschewed such words because of the suspicion that they register differences primarily of subject matter or content, rather than authorship. Whereas to differentiate authors, John Burrows used the 150 most frequent words, perforce for the most part ‘function words’,2 Hoover found that the 600 or 700 most frequent words (‘at which point almost all the words are content words’3) enhanced discrimination with Burrows's method. Hoover's discovery complements Cyril and Dominique Labbé's method of intertextual distances, which makes use of all the available words involved in textual comparisons.4 An experiment employs words contained in four comedies of Shakespeare (All's Well That Ends Well, and/or As You Like It, and/or Much Ado About Nothing, and/or Twelfth Night), preferably occurring in all four plays, as well as in Henry VIII (All Is True), and Fletcher's The Woman's Prize, and/or The Island Princess, and/or Demetrius and Enanthe, and/or Valentinian, and/or The Loyal Subject, and/or Monsieur Thomas, and/or The Chances, preferably occurring in all seven plays. All the selected words must occur in the anchor play, Henry VIII. An assumption was made that words in Henry VIII were either Fletcher-favouring or Shakespeare-favouring. A list of words which occur more on average in the Fletcher plays than in the Shakespeare F1 comedies (Group1) was compiled, as well as the reverse (Group 2).

Read the paper · More papers on PaperTik