Performance Analysis of Vision Transformer Based Architecture for Cursive Handwritten Text Recognition

Avishek Chowdhury, Md. Ahasan Hossen, Mohammad Shahidul Hasan, Sheikh Md Rukunuddin Osmani · 2023

Computers have been given vision by researchers across the world for many years. Now it is the era of digitization. Recognizing handwritten text is a must for a computer vision system. Due to the variation and complexity of the cursive writing style, the holistic approach is mostly used for the recognition of cursive scripts. Though Convolutional Neural Network (CNN) based models have been employed in literature for Holistic Handwritten Text Recognition (HTR) of different domains, recent breakthroughs in the image classification Vision Transformer (ViT) based models have not been utilized for HTR so far. In this research, we have designed a ViT-based model for the HTR of various cursive scripts. To validate the performance of the model, various handwritten datasets of cursive scripts have been used. The notable finding includes that the accuracy of the designed model has increased by up to 26% after applying the image data augmentation techniques.

Read the paper · More papers on PaperTik