Scribe versus authorship attribution and clustering in historic Czech manuscripts: a case study with visual and linguistic features
Aleksej Tikhonov, Klaus M. Müller · Digital Scholarship in the Humanities · 2020
Abstract For the identification of scribes and authors in handwritten documents, methods from classical linguistic analysis are combined with modern computer vision approaches to enhance the knowledge discovery process. One important finding is that it is possible to train neural networks for automatic transcription of handwritten documents and to use these transcriptions as input for statistical analysis. Furthermore, hypotheses about scribes can be tested by extracting visual handwriting features and clustering them. From a linguistic point of view, the R package stylo is a useful tool to analyse and cluster texts. Unfortunately, it only achieves a high level of accuracy with longer texts. For texts under 5000 words it is more suitable to measure their Euclidean distance based on a set of linguistic features. Both approaches, the analysis with stylo and the Euclidean distance, in combination with neural networks for automatic transcription and clustering allow for more precise statements about the relationship between texts, authors and scribes, even if the documents are under 1,000 words.