An SMT Approach to Automatic Annotation of Historical Texts
Eva Pettersson, Beáta Megyesi, Jörg Tiedemann · 2013
In this paper we propose an approach to tagging and parsing of historical text, using characterbasedSMT methods for translating the historical spelling to a modern spelling before applyingthe NLP tools. This way, existing modern taggers and parsers may be used to analyse historicaltext instead of training new tools specialised in historical language, which might be hardconsidering the lack of linguistically annotated historical corpora. We show that our approachto spelling normalisation is successful even with small amounts of training data, and thatit is generalisable to several languages. For the two languages presented in this paper, theproportion of tokens with a spelling identical to the modern gold standard spelling increasesfrom 64.8% to 83.9%, and from 64.6% to 92.3% respectively, which has a positive impact onsubsequent tagging and parsing using modern tools.