Preprocessing and structural feature extraction for a multi-fonts Arabic/Persian OCR
Mandana Kavianifar, Adnan Amin · 1999
English and Chinese are languages which have attracted tremendous interest from character recognition researchers. In contrast, research in the field of character recognition for Arabic/Persian scripts faces major problems, mainly related to their unique characteristics, like being cursive, the multiple shapes of one character in different positions in a word, and the connectivity of characters on the baseline. The work proposed in this paper consists of three major phases. After digitizing the text, the original image is transformed into a gray-scale image using a 300-dpi scanner. Different pre-processing steps are then applied to the image file. In the next phase, sub-words of all words are recognized and global features for each word are extracted. Contour tracing plays a very important role in the feature extraction phase.