Image Pre-processing on NumtaDB for Bengali Handwritten Digit Recognition
Ovi Paul · 2018
NumtaDB is by far the largest data-set collection for handwritten digits in Bengali language. This is a diverse dataset containing more than 85000 images. But this diversity also makes this dataset very difficult to work with. The goal of this paper is to find the benchmark for pre-processed images which gives good accuracy on any machine learning models. The reason being, there is no available pre-processed data for Bengali digit recognition to work with like the MNIST for English digits.