Sankhya: An Unbiased Benchmark for Bangla Handwritten Digits Recognition

A. B. M. Alim Al Islam, Fuad Rahman, AKM Shahariar Azad Rabby · 2019

The rise of artificial intelligence technology along with machine and deep learning are opening up almost limitless possibilities. In recent years, application-based researchers in machine learning and deep learning have started developing solutions for many practical problems. Handwriting recognition is one such area of interest. Bangla, being the seventh most spoken language in the world, is not an exception. However, unlike English, there have not been concerted formal attempts in building a benchmark in comparing the different approaches reported in the literature, mainly because of the lack of openly and freely available datasets and diversity of the approaches without formal comparative studies. In this research paper, we seek to rectify this gap. We have focused on benchmarking five robust algorithms: K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Random Forest (RF), Multi-layer Perceptron (MLP), Convolutional Neural Network (CNN), on all publicly available Bangla handwriting digits datasets, including Ekush, NumtaDB, CMARTdb, and BDRW. NumtaDB itself is a collection of five handwriting datasets. We have worked on fine-tuning these algorithms by finding the best possible hyper-parameters of these algorithms. It is our hope that Sankhya will work as a beginning point of an open and verifiable benchmarking process that we plan to repeat every two years for now on and set a standard for testing and validating newer and novel algorithms that will be reported in this area in the future. In addition, we have extensively compared our research with other states of the art research and our versions of these algorithms are now outperforming every reported result on these datasets.All the datasets we used are open-sourced. In addition, we are making the.csv version of these datasets available in public GitHub. Of all the models we tested, the Sankhya CNN model performed the best for all these datasets, which we fine-tuned specifically for Bangla character recognition. We are making this CNN model available in public GitHub.

Read the paper · More papers on PaperTik