Impact of feature selection and engineering in the classification of handwritten text
Anupama Kaushik, Himanshu Gupta, Digvijay Singh Latwal · International Conference on Computing for Sustainable Global Development · 2016
Feature selection forms an important aspect of machine learning and character recognition. It is a process of selecting the most important features (attributes) from the dataset. Accurate feature selection results in significant reduction in the number of irrelevant (constant, redundant) attributes thereby, reducing the processing time and increasing the accuracy of the model without any loss of information. This paper focuses on the impact of feature selection and engineering in the classification of handwritten text by identifying and extracting those attributes of the training dataset that will contribute most towards the classification task using classifiers like J48, NaiveBayes and Sequential Minimal Optimization (SMO). This results in improved accuracy of the classifiers as compared to the work reported earlier. Further, a comparative performance evaluation of the classifiers used for OCR and pattern recognition is done. Initial classification performance of all the classifiers listed above was recorded on the raw dataset. Finally, the dataset was transformed after performing relevant feature selection and engineering on its attributes. The same classifiers were again trained on the transformed dataset and their accuracy was recorded. This paper uses the widely used MNIST dataset of handwritten digits for training the classifiers.