GTReZ - Gujarati Text Recognition using Zone Segmentation

Joel D’Silva, Aryan Koul, Madhur Thakkar, Neha Bharambe, Yashika Kuckian, Archana Shirke, S. Dugad · 2023

Many ancient books and parchments are available in various parts of India which are deteriorating with time. Preserving these manuscripts and writings is equal to preserving the heritage of our country. Digitization increases the efficiency and protects the records no matter what natural disaster, robbery, or loss happens. According to the statistical data collected by the National Mission for Manuscripts, around 30% of the manuscripts available in India are written in Gujarati language. The proposed project represents an effective, robust, and highly accurate system to carry out this task with convenience. It consists of a model that will convert the typed newspaper text into online machine format that can be edited easily. The traditional approaches required each and every combination of consonants and modifiers to be considered. This would result into a large number of classes [47*12 = 564]. Therefore, an approach is implemented to reduce the number of classes where the text line is divided into three parts horizontally on the basis of word alignment. The words are divided into vowel diacritic above the base letter, the base letter and the vowel diacritic below the base letter. After the post-processing is done, the final output will be the combination of the three zones. This will result in drastic reduction in the number of classes involved in training the model from 564 to 79. Along with inference, the proposed model will also be able to detect text blocks from columns of text and will be highly efficient for Gujarati text mainly focusing on vowels.

Read the paper · More papers on PaperTik