Preparing Persian/Arabic Scanned Images for OCR
Sajad Shirali-Shahreza, Mohammad Taghi Manzuri, Sajad Shirali-Shahreza · 2006
Digital documents are widely used today. So converting written documents such as books to digital documents is unavoidable. The most popular method for doing this is OCR. Usually documents are scanned and then scanned images are sent to OCR. Scanned images need some preprocessing in order to be used in OCR efficiently. In this paper, we introduce a method for preparing scanned Persian/Arabic printed texts for OCR. Our method considered especial features of Persian/Arabic scripts such as dots and connecting characters. Main phases of our work are converting grayscale image to binary image, removing straight lines and frames and identifying picture components