Normalizing Medieval German Texts: from rules to deep learning

Natalia Korchagina · Zurich Open Repository and Archive (University of Zurich) · 2017

The application of NLP tools to historical texts is complicated by a high level of spelling variation.Different methods of historical text normalization have been proposed.In this comparative evaluation I test the following three approaches to text canonicalization on historical German texts from 15 th -16 th centuries: rule-based, statistical machine translation, and neural machine translation.Character based neural machine translation, not being previously tested for the task of normalization, showed the best results.

Read the paper · More papers on PaperTik