An approach to unsupervised historical text normalisation
Petar Mitankin, Stefan Gerdjikov, Stoyan Mihov · 2014
We present a novel approach to unsupervised noisy text correction. Our approach is based on automatic extraction of historical variation patterns by analysing the structure of the words from a historical corpus and comparing it with the structure of the contemporary dictionary. Based on the extracted variation patterns the core candidate generator, REBELS, produces correction candidates even outside the modern dictionary. Further, the sentence correction is complemented with a modern language model combined in a log-linear model. The quality of our unsupervised approach is empirically compared against a supervised system competitive with the state-of-the-art supervised text normalisation systems. The experiments show that our system delivers 81.79% normalisation accuracy of 17th century English historical texts in a fully unsupervised setup.