Benchmark dataset for offline handwritten character recognition
Adeel Yousaf, Muhammad Jaleed Khan, Muhammad Ali Imran, Khurram Khurshid · 2017
A benchmark dataset is the first and foremost step in Handwritten Character Recognition (HCR). This paper provides comprehensive detail about a newly compiled dataset known as Handwritten Characters Dataset (HCD), applied for recognition of handwritten English characters (A-Z) and digits (0–9). Compilation of this dataset is of great significance as it contains segmented characters rather than words, sentences or paragraphs. Total 150 writers belonging to different ages, gender and professional backgrounds contributed in this dataset. HCD is publically available free of cost1. HCD can be used for training of classifiers like Neural Network (NN), Hidden Markov Model (HMM) or Support Vector Machine (SVM) etc. which further can be used for a wide variety of applications of handwritten text recognition. Experimental analysis of HCD has depicted very promising results compared to state of art datasets.