Fine-Tuning an Arabic OCR Model using Tesseract 5.0

Omar Samir, Yousef Waleed, Ibrahim Ahmed, Omar El-Gohary, Shaden El-Naggar, Farida Fawzy, Malak Emam, Mayar Emam, Mohamed Taher Alrefaie · 2024

Optical Character Recognition (OCR) is a crucial technology for the digital processing and preservation of textual information. While significant progress has been made in OCR for commonly used languages, the recognition of Arabic script poses unique challenges due to its complex linguistic characteristics. This research paper presents a comprehensive approach to collecting data and training an Arabic OCR system using the latest version of the Tesseract OCR engine. Integrating the customized Tesseract model, the study achieved high accuracy in text recognition, as measured by character error rate (CER) and word error rate (WER). The findings highlight the complexities of the Arabic language, including diacritics and ligatures, and discuss the limitations of the current approach. This study contributes to the advancement of optical character recognition, providing a robust Arabic OCR solution that can be leveraged in various applications, from digitizing historical documents to automating text extraction.

Read the paper · More papers on PaperTik