A New Benchmark Dataset for Handwritten Character Recognition
Laurens van der Maaten · 2009
The report presents a new dataset of more than 40 , 000 handwritten characters. The creation of the new dataset is motivated by the ceiling effect that hampers experiments on popular handwritten digits datasets, such as the MNIST dataset and the USPS dataset. Next to a character labeling, the dataset also contains labels for the 250 writers that wrote the handwritten character, which gives the dataset the additional potential to be used in forensic applications. The report discusses that data gathering process, as well as the preprocessing and normalization of the data. In addition, the report presents the results of initial classification and visualization experiments on the new dataset, in an attempt to provide a base line for the performance of learning techniques on the dataset.