ALNASIKH: An Arabic OCR System Based on Transformers

Ahmad Mortadi, Ahmed Mohamed, Ahmed Talima, Ahmed Alkhattip, Ahmed K. Ibrahim, Ahmed Mohamed Elhassan Elfaki Osman, Yasser Hifny · 2023

Recognizing Arabic text from images is a challenging task due to the intricate nature of the Arabic script, which incorporates a large number of characters and intricate diacrit-ical marks. The unique characteristics of the script, including ligatures and overlapping characters, add complexity to the OCR process. To address these challenges, we propose utilizing the TrOCR architecture, which employs deep learning techniques, specifically transformers, for text recognition tasks and has shown superior performance in text recognition tasks. In this research, we evaluate the effectiveness of the TrOCR model architecture for recognizing Arabic text by proposing a setup for its encoder and decoder tailored specifically for Arabic script. Our proposed setup takes into account the unique characteristics and complexities of Arabic text, allowing the TrOCR model to effectively capture and interpret the intricacies of the script. Through extensive experimentation and evaluation, our results demonstrate the efficacy of the proposed approach. When applied to the task of recognizing text from scanned Arabic documents at the word-level, our approach achieves remarkable accuracy. The average character error rate (CER) is measured at 0.8, indicating a high level of precision in recognizing individual characters. Similarly, the word error rate (WER) is measured at 2.3, indicating accurate transcription of complete words. These results highlight the potential of the TrOCR model architecture, in conjunction with our proposed setup, to significantly improve the accuracy of OCR systems for Arabic text recognition.

Read the paper · More papers on PaperTik