Comparative Analysis of Handwritten Digit Recognition using MLP, CNN, LeNet-5 Model’s
Mughesh Kumar N R, Siddharth Joshi, Pranava Preethivardhan Chanduri, S. Sivakumar · 2024
The analysis of three distinct neural network architectures for Handwritten Digit Recognition (HDR): Multi-Layer Perceptron (MLP), Simple Convolutional Neural Network (CNN), and LeNet-5. The MNIST dataset, an internationally renowned benchmark made up of 70,000 grayscale photos of handwritten numbers, is used to assess the models. To alleviate overfitting and improve model resilience, the data is pre-processed and split into training and validation sets using K-Fold cross-validation. The MLP architecture employs a fully connected layer (FCL) structure with three layers, achieving a training accuracy of 87.12% and a validation accuracy of 95.78% in 115 seconds of training time. The Simple CNN consists of two convolutional layers (CL) and two fully connected layers, yielding, after 395 seconds of training, a validation accuracy of 99.06% and a training accuracy of 99.76%. LeNet-5, with two convolutional layers and three fully connected layers, takes 168 seconds to train and achieves 99.38% training accuracy and 98.89% validation accuracy. Performance evaluation reveals that while the Simple CNN exhibits the highest accuracy (99.15%) on the test set, LeNet-5 offers a more balanced trade-off between training time and accuracy. This work underscores the varying capabilities of MLP, Simple CNN, and LeNet-5 architectures in HDR tasks, providing insights for selecting the appropriate model based on specific application requirements.