Comparing Rule-based and SMT-based Spelling Normalisation for English Historical Texts

Gerold Schneider, Eva Pettersson, Michael Percillier · Zurich Open Repository and Archive (University of Zurich) · 2017

To be able to use existing natural language processing tools for analysing historical text, an important preprocessing step is spelling normalisation, converting the original spelling to present-day spelling, before applying tools such as taggers and parsers. In this paper, we compare a probablistic, language-independent approach to spelling normalisation based on statistical machine translation (SMT) techniques, to a rule-based system combining dictionary lookup with rules and non-probabilistic weights. The rule-based system reaches the best accuracy, up to 94% precision at 74% recall, while the SMT system improves each tested period.

Read the paper · More papers on PaperTik