Automatic Handwriting Recognition in Historical Documents
Andreas Fischer · Series in machine perception and artificial intelligence · 2020
Automated reading of handwriting is an intriguing and largely unsolved problem in computer science [Lorette (1999)].Despite decades of continuous progress, the goal to match or even surpass the astounding capabilities of humans to read handwritten text is still far from being reached [Bunke and Varga (2007)].One of the main reasons is that automatic systems still lack semantic understanding that is often needed to decipher difficult cases, such as handwritten notes taken years ago with crossed out text and abbreviated words.This is especially true for ancient manuscripts, where the transcription of a phrase may only be found through expert knowledge in paleography, linguistics, and history.Standard recognition techniques for optical character recognition (OCR) that are based on character segmentation, followed by character recognition, typically fail in the case of handwriting.When facing variable character shapes and touching characters, recognition is needed to perform segmentation and, vice versa, segmentation is needed to perform recognition.This conundrum, also known as Sayre's paradox [Sayre (1973)], requires special recognition techniques that perform segmentation and recognition at the same time.A well-established approach to handwriting recognition is based on Hidden Markov Models (HMM) [Ploetz and Fink (2009)], inspired by their success for modeling sequences in speech recognition [Rabiner (1989)].In the past