Script Identification of Central Asia Based on Fused Texture Features

Xing-kun Han, Alimjan Aysa, Hornisa Mamat, Nurbiya Yadikar, Kurban Ubul · 2018

Script identification is an important step in multi-script recognition. Despite the achieved results in this field, the identification of Central Asian scripts has not been considered in-depth. In the Central Asian region, there are many similar scripts, and the traditional texture features can not discriminate them accurately. This paper proposes a script identification method based on fused texture features for Central Asian document images. On preprocessed multilingual document images, the method first performs Non-subsampled Contourlet Transform (NSCT), and then extracts Tamura texture features of the generated sub-bands. A Support Vector Machine (SVM) classifier is trained for classification. For experimental evaluation, it is collected a dataset of 30, 000 document images for 10 scripts, such as Arabic, Chinese, English, Russian, Kazakhstan, Turkish, Uyghur, Kyrgyzstan, Mongolian and Tibetan. The experimental results show that the proposed method can extract multi-scale and multi-directional texture features, and the fusion of texture features leads to superior performance of script identification.

Read the paper · More papers on PaperTik