A Novel Minimal Script for Arabic Text Recognition Databases and Benchmarks

Husni Al-Muhtaseb, Sabri A. Mahmoud, Rami S. Qahwahi · 2009

This paper presents a minimal Arabic text that covers the different basic shapes of Arabic alphabet (viz. standalone, initial, medial, and terminal). It is designed with minimal repetition of character shapes in the minimal text. The novelty of the suggested script could be seen from different perspectives. It enables the collection of handwritten text from different writers with minimized effort and time. It is enough for a writer to write three lines of meaningful Arabic text to cover all possible character shapes, a total of 125 shapes. The written text is designed to have even distribution of letter frequencies. This assures enough samples of all character shapes when text is collected from enough number of writers. The same is true for printed Arabic text. This is especially useful when using large number of features with classifiers that require large number of samples for each category. Hidden Markov Models and Neural networks are two examples of these classifiers. The use of the minimal text enables proper training, as all Arabic character shapes are present with adequate frequency, hence resulting in higher recognition rates. This is not the case with natural text where the frequency of some Arabic characters differ widely, where in some cases 100 folds or more. The proposed minimal text may be used to build a data base of handwritten Arabic text collected of many writers. This covers the need for a database in the research of Arabic handwritten text recognition and benchmarking. In addition, this paper presents statistical analysis of Arabic corpora for estimating the number of occurrences of the different shapes of Arabic characters in large corpora. The frequency of Arabic characters could be used in different applications. In this research work, it was utilized in enhancing the search for the minimal Arabic text.

Read the paper · More papers on PaperTik