Bigrams of Syntactic Labels for Authorship Discrimination of Short Texts

Graeme Hirst, Olga Feiguina · Literary and Linguistic Computing · 2007

We present a method for authorship discrimination that is based on the frequency of bigrams of syntactic labels that arise from partial parsing of the text. We show that this method, alone or combined with other classification features, achieves a high accuracy on discrimination of the work of Anne and Charlotte Brontë, which is very difficult to do by traditional methods. Moreover, high accuracies are achieved even on fragments of text little more than 200 words long.

Read the paper · More papers on PaperTik