A Multilingual Handwritten Character Dataset: T-H-E Dataset
Gaye Ediboglu Bartos, Yaşar Hoşcan, András Kauer, Éva Hajnal · Acta Polytechnica Hungarica · 2020
The absence of handwritten special Latin character datasets prompted the creation of the T-H-E Dataset (Turkish-Hungarian-English handwritten character dataset) contributing to the recognition of multilingual handwritten texts.This paper represents a public-domain dataset including handwritten Turkish, Hungarian and English characters collected from 200 participants.The T-H-E Dataset is formed from 78 different letters represented in 156000 binary characters including both the upper and lower-case versions.The dataset can be downloaded from the web in six different versions enabling users to combine the different alphabets for different recognition purposes.The evaluation of the dataset is carried out by applying the same deep learning architecture on the T-H-E dataset and the EMNIST dataset.The dataset is publicly available at https://github.com/bartosgaye/thedataset.