Odia Handwritten Character Recognition with Noise using Machine Learning
Anupama Sahu, Sumita Mishra · 2020 IEEE International Symposium on Sustainable Energy, Signal Processing and Cyber Security (iSSSC) · 2020
Optical Character Recognition (OCR) is a burning technology to recognize text inside images in the current era, such as: scanned documents and photos. OCR technology is used to convert virtually any kind of images containing written text (handwritten or printed) into machine-readable text data. Several research work have been done on the recognition of different foreign languages such as Chinese, English and Japanese Scripts. In India, there are 22 official languages such; Marahati, Punjabi, angala, Odia, etc. Odia is one of the majorly spoken languages of Odisha, a premier Eastern state of India. Many Indian scripts languages have been researched and yielded results of good accuracy rate such as Devanagari, Telugu scripts. There is a need for research on the languages of the Eastern part of the country such as Odia. In this paper it has been implemented for data preprocessing and classification model for offline odia handwritten character with and without noise. Hence this research work has been strived towards buildout of a narrative machine learning algorithm for classification of Offline Odia handwritten Character using Naive Bayes and Decision Table in Waikato Environment for Knowledge Analysis (WEKA) environment. It has been observed noiseless character is better than the noise character in both classification techniques such as: Naive Bayes and Decision Table.