Automatic language identification of bilingual English and Farsi scripts
Hamideh Rezaee, Masoud Geravanchizadeh, Farbod Razzazi · 2009
In general, printed documents may contain several different languages. Therefore, to use Optical Character Recognition (OCR) for multi-lingual documents, it is necessary to automatically separate these languages. In this paper, we describe a method for identification of printed Farsi and English text from images of documents in line and word levels. The proposed algorithm is developed based on statistical and shape-based features. The accuracy of this method is around 96.05%.