Tesseract OCR Recognition Based on Arabic Machine-Printed Document
Rakesh Jagdish Ramteke, Mohammed Rashed Ali Omar Al Maamari · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2023
This paper provides technical aspects and the context of Recognizing and Detecting Arabic characters using Tesseract OCR Engine.OCR engine is freely available and gives a better result and also is supporting many languages such as Arabic etc.The procedure begins by transforming the Arabic documents into machine format (scanning) and then recognizing as well as extracting the text using the PyTesseract library.The OCR is a system that can afford the considerable values of split errors, particularly while working with cursive languages like the Arabic language with repeated overlapping between letters.Moreover, The performance is 99.5 accuracy in OCR-tesseract for converting the Arabic image documents to text editable.