Comparison Study of Machine Learning Techniques for Letter Recognition

Rizal Dwi Prayogo, Siti Amatullah Karimah · 2022

Classification is a supervised learning technique that can learn to assign a collection of attributes to the classes. This paper aims to recognize the 26 uppercase letters of the English alphabet using classification algorithms. The objective of this study is to discover the best classifier for letter recognition which has many practical applications for the improvement of the automation process such as document reading, word processing, or mail sorting. The letter recognition dataset consists of 20000 samples with 16 extracted attributes that represent alphabet letter, font type, aspect ratio, linear magnification, and vertical and horizontal position to identify the 26 class values of uppercase letters from A to Z. An information-based attribute selection and multiple classifiers are proposed involving Bayesian classifiers, functions classifiers, lazy classifiers, and tree classifiers with multi-test options to compare the recognition performance. The classification performances are evaluated by the measures of accuracy, precision, recall, F-measure, the Receiver Operating Characteristics area, the Matthews Correlation Coefficient, Root Mean Squared Error, and time to build a model. The results show that the highest average accuracy of the test option is given by 80%-split, 10-fold Cross-validation, 90%-split, and 5-fold Cross-validation, respectively. The overall study shows that Random Forest is the best classifier in the letter recognition dataset for most measurements, except for the processing time.

Read the paper · More papers on PaperTik