Normalizing Medieval German Texts: from rules to deep learning
Natalia Korchagina · Zurich Open Repository and Archive (University of Zurich) · 2017
The application of NLP tools to historical texts is complicated by a high level of spelling variation.Different methods of historical text normalization have been proposed.In this comparative evaluation I test the following three approaches to text canonicalization on historical German texts from 15 th -16 th centuries: rule-based, statistical machine translation, and neural machine translation.Character based neural machine translation, not being previously tested for the task of normalization, showed the best results.