Efficient Authorship Attribution Method Using Ensemble Models Built by Knowledge Distillation

Wataru Sakurai, Masato Asano, Daisuke Imoto, Masakatsu Honma, Kenji Kurosawa · 2023

Authorship attribution is a task to identify the author of given documents. It is often treated as a classification task that predicts the author of a written text among given candidates. From practical application perspectives, such as forensic science, the interpretation of results is essential, and the ability to perform authorship attribution with many sentences at high speed will become increasingly important for developing authorship attribution methods. Recently, transformer-based large-scale language models have been developed in natural language processing, and high accuracy has been achieved in various fields. The attention mechanism implemented in the encoder part of the transformer is called self-attention, which calculates the degree of attention given to each word in a sentence. However, in classification problems such as author attribution, the only evidence used as a basis for judgment is the attention distribution, which is a single dimension of the matrix calculated by self-attention and is computationally inefficient. In addition, since this attention distribution is obtained for the number of attentions in the model, its contribution to estimation results and utilization is difficult to examine. In this study, we propose an authorship attribution method that ensembles compact models that substitute each attention of a fine-tuned transformer made by knowledge distillation. We also examine the contribution of each attention to the estimation results based on the accuracy of each ensemble model. In authorship attribution experiments using the works of modern Japanese authors, the ensemble of distillation models achieved results comparable to those of bidirectional encoder representations from transformers. Furthermore, if the process can be ideally parallelized, the computational time required for inference by the proposed method can be significantly reduced.

Read the paper · More papers on PaperTik