Digital Peter: New Dataset, Competition and Handwriting Recognition Methods

Mark Potanin, Denis Dimitrov, Alex Shonenkov, Vladimir Bataev, Denis Konstantinovich Karachev, Maxim Novopoltsev, Andrey V. Chertok · 2021

This paper presents a new dataset of Peter the Great’s manuscripts and describes a segmentation procedure that converts initial images of documents into lines. This new dataset may be useful for researchers to train handwriting text recognition models as a benchmark when comparing different models. It consists of 9694 images and text files corresponding to different lines in historical documents. The open machine learning competition ”Digital Peter” was held based on the considered dataset. The baseline solution for this competition and advanced methods on handwritten text recognition are described in the article. The full dataset and all codes are publicly available.

Read the paper · More papers on PaperTik