A new algorithm for Arabic optical character recognition

Omar Al‐Jarrah, Samer Al-Kiswany, Bashar Al-Gharaibeh, Mohamed Fraiwan, Hani Khasawneh · International Conference on Artificial Intelligence · 2006

In this paper we present a new approach for designing an Optical Character Recognition (OCR) system for Arabic Alphabet. Our approach addresses mere text images; using customized techniques. We have divided the problem at hand into three phases: preprocessing and line extraction, segmentation, and character recognition. Extensive histogram and thresh-holding were considered in line extraction, image filtering and noise reduction. Main line algorithm was proposed for segmentation, and a neural network was used for character recognition. The segmentation algorithm is based on a new and novel idea that exploits the nature of the Arabic language. It introduces the concept of the main line to tokenize the text, and it generates a set of 33 different tokens that represent the 28 Arabic characters and their different shapes and variation. The system generates tokens from the text and compares them with the reference tokens. Experimental results illustrate the effectiveness of the system using different font types and sizes with an average recognition rate of 87%.

Read the paper · More papers on PaperTik