PHCWT: A Persian Handwritten Characters, Words and Text dataset
Mojtaba Montazeri, Rasoul Kiani · 2017
Nowadays, different languages involve certain datasets to recognize optical characters in handwritten materials both online and offline. Although the valuable datasets in Persian language have been approved, there are several deficiencies such as insufficient characters, words and texts in a single database. This paper proposed a new dataset Persian Handwritten Characters Words Text (PHCWT). This unique dataset is the first extensive and integrated collection of Persian handwritten characters, words and texts, which can be used in various fields such as offline handwritten optical character recognition, word segmentation, text extraction, baseline detection, author detection, etc. This dataset contains a total of 51,200 character images, 3,600 word images and 400 texts collected by 400 contributors. One of the most important features of PHCWT is the adoption of skeleton and edge as two techniques for image extraction. Each contributor was handed 3 forms containing the characters in the Persian alphabet (in 4 modes: single, beginning, middle and ending), 9 words matching polymorph characters and an engineered text including all Persian characters. The PHCWT is available to researchers.