Authorship Attribution in Greek Literature Using Word Adjacencies
Rizos-Theodoros Chadoulis, Andreas Nikolaou, Constantine L. Kotropoulos · 2022
Authorship attribution stems from the idea that one can use a text to derive useful conclusions about its author. It is a rather old idea, that has recently gained momentum due to the major technological leaps achieved in computer science. These leaps enabled researchers to process large amounts of information in reasonable time and to employ sophisticated algorithms for extracting textual features that may indicate something useful about the author. In this paper, a method for authorship attribution is investigated that resorts to Word Adjacency Networks (WANs). The method was originally proposed for author attribution in English corpora by Segarra, Eisen, and Ribeiro [23]. The paper builds upon this method by demonstrating its capability to capture the stylometric patterns of authors in Greek literature. Contrary to English, Greek is a strongly synthetic language. The data used in the experiments comprise literature pieces from 20 Greek authors that belong to two different groups. The first 9 (i.e., Eftaliotis, Karkavitsas, Kondylakes, Moraitides, Nirvanas, Papadiamantes, Rhoides, Vikelas, and Viziinos) form the first group. The remaining 11 (i.e., Delta, Empirikos, Karagatsis, Kastanakis, Kontoglou, Mirivilis, Politis, Prevelakis, Terzakis, Theotokas, and Venezis) form the second group whose nucleus is the so-called Generation of 30s’. The problem is formulated mathematically and the attribution algorithm is tested for a wide range of parameter values. The experimental findings show that authorship attribution with WANs can give satisfactory results as long as there are adequate training data. After fine tuning the parameters of the method, the total attribution accuracy was found to be in the first group of authors and in the second group of authors. This demonstrates the potential of the method and its applicability to morphologically and syntactically rich languages.