A web-based tool for segmentation and automatic transcription of historical documents

Fouad Slimane, Andrea Mazzei, Orlin Topalov, Greta Verzi, Frédé́ric Kaplan · 2017

This paper describes a web-based system for page segmentation and text recognition of historical documents. The system is organised following a pipeline of 4 steps : 1) digitisation, 2) preprocessing, 3) textline extraction, and 4) handwritten text recognition based on hidden Markov models. In this study we used to evaluate the system the “Statuti del Doge Tiepolo”, a 14thcentury manuscript written in Gothic script, digitized and transcribed by an expert paleographer, The paper discusses the initial performances of interface and processing pipeline in this context.

Read the paper · More papers on PaperTik