Stacked authorship attribution of digital texts

José Eleandro Custódio, Ivandré Paraboni · Expert Systems with Applications · 2021

In computational authorship attribution (AA) – the task of identifying the author of a given text based on a set of possible candidates – existing differences across domains, languages or input settings may require using knowledge from multiple sources, ranging from surface character patterns to deeper semantics and others. Moreover, since increasing the model complexity may easily lead to overfitting, sources of this kind have to be selected judiciously according to each particular input. Based on these observations, this article introduces a novel approach to AA consisting of stacked classifiers built from multiple knowledge sources - words, characters, part-of-speech n-grams, syntactic dependencies, word embeddings and more - that are dynamically included in the AA model according to the relevant input. In doing so, we would like to show that a stacking approach not only outperforms previous work in the field, but also that dynamic model selection outperforms the use of any of the individual components alone. The current model - called DynAA - is evaluated in a number of AA scenarios covering multiple languages, domains and input sizes, and is shown to generally outperform a number of baseline alternatives, including convolutional neural networks, BERT and others.

Read the paper · More papers on PaperTik