1 Treebanks for historical languages and scalability

Marius L. Jøhndal · 2020

Treebanks for historical languages and scalability 1 IntroductionHistorical linguistics, whether synchronic or diachronic, is by definition based on corpora.Since we do not have access to the intuitions of native speakers we can only test linguistic hypotheses about historical languages by systematically collating information from our corpus of texts.For questions that typically concern linguists, this often means identifying every occurrence of a particular phenomenon in the corpus, analysing, classifying and counting the occurrences and then using this for testing hypotheses about the structure of the language.This can be done manually, but this is time-consuming and error-prone.As Haug (2015) points out, while reading the text and manually collating information from it is essential for hypothesis formation it is much less useful for hypothesis testing.Even if the text is in electronic form, it is easy to overlook an example, record it incorrectly or fail to apply test criteria consistently over time.This paper focuses on treebanks, which are corpora that have been annotated with morphosyntactic information so that we can extract linguistic structures like 'verb with an accusative noun'.High-quality treebanks for a range of historical languages now exist and are widely used in historical linguistic research.This includes treebanks that follow the Penn-style of annotation, e.g. the Penn-Helsinki Parsed Corpus of Middle English (Kroch and Taylor 2000), the Penn-Helsinki Parsed Corpus of Early Modern English (Kroch, Santorini, and Delfs 2004), the Penn-Helsinki Parsed Corpus of Modern British English (Kroch, Santorini, and Diertani 2016), the Tycho Brahe Parsed Corpus of Historical Portuguese (Galves and Britto 2002) and the Icelandic Parsed Historical Corpus (Wallenberg et al. 2011), as well as dependency-based treebanks, e.g. the Index Thomisticus (Passarotti 2007), the Ancient Greek and Latin Dependency Treebanks (Bamman and Crane 2011; Celano, Crane and Almas 2014), the PROIEL Treebank (Haug and Jøhndal 2008, Haug, Eckhoff et al. 2009), the ISWOC Treebank (Bech and Eide 2014) and the TOROT Treebank (Eckhoff and Berdičevskis 2015).A key challenge in building treebanks for historical languages is lack of resources.Funding is limited and there are few existing computational language resources like taggers and parsers available.At the same time, the task is complex and experts on the language have to devote a significant amount of time to Open Access.

Read the paper · More papers on PaperTik