HWR200: New open access dataset of handwritten texts images in Russian
Antiplagiat, I. O. Potyashin, Mariam Kaprielova, Yu. V. Chekhovich, A. S. Kildyakov, Temirlan Seil, Evgeny Finogeev, Andrey Grabovoy · Computational Linguistics and Intellectual Technologies · 2023
Handwritten text image datasets are highly useful for solving many problems using machine learning. Such problems include recognition of handwritten characters and handwriting, visual question answering, near-duplicate detection, search for text reuse in handwriting and many auxiliary tasks: highlighting lines, words, other objects in the text. The paper presents new dataset of handwritten texts images in Russian created by 200 writers with different handwriting and photographed in different environment1 . We described the procedure for creating this dataset and the requirements that were set for the texts and photos. The experiments with the baseline solution on fraud search and text reuse search problems showed results of results of 60% and 83% recall respectively and 5% and 2% false positive rate respectively on the dataset.