Bengali Handwritten Grapheme Recognition Using CutMix-Based Data Augmentation
Zhenrong Zhang · 2021
Recent work has demonstrated that deep learning has been widely used in various challenging images classification tasks. In our paper, we describe Bengali handwritten grapheme classification based on deep learning. We take advantage of a deep learning architecture called convolutional neural networks (CNNs) due to its splendid performance, effectiveness of convolutional operations on many different visual recognition tasks. In addition, we propose Cutmix-based data augmentation to virtually enlarge the training dataset size and avoid over-fitting in performing the further recognition tasks. Comparisons have been made between Cutmix-based data augmentation and other data augmentation techniques (such as flipping, distorting, rotation, mixup) on four widely-used CNN architectures: Inception, ResNet50, DenseNet169 and EfficientNet B1. During experiments, CutMix consistently outperforms the other augmentation strategies on Bengali handwritten grapheme classification tasks, by making efficient use of training pixels and retaining the regularization effect of regional dropout. Unlike previous augmentation methods, our CutMix-trained EfficientNet classifier results in consistent performance gains in Bengali handwritten grapheme detection and handwritten image captioning benchmarks. Moreover, CutMix improves the model robustness against input corruptions and its out-of-distribution detection performances. Therefore, CutMix proved to be a naturally enhanced augmentation strategy with superior concision and effectiveness in classifying Bengali handwritten grapheme.