Automatic recognition of common Arabic handwritten words based on OCR and N-GRAMS

Laslo Dinges, Ayoub K. Al-Hamadi, Moftah Elzobi, Andreas Nürnberger · 2017

Comprehensive databases are vital for training and validation of word recognition systems. To overcome the lack of offline databases of Arabic handwritten words, especially regarding the generality of the underlying vocabulary, we used a synthesis system to generate a database of common Arabic handwritings. Subsequently, we validate a new word recognition system on these synthetic handwritings, to analyze the performance of its segmentation, character recognition, and error correction module. We found, that a dynamic character classifier, that is capable to adapted to the variations that are caused by the segmentation, clearly improves word recognition accuracy. For error detection and correction, n-grams as well as the Levenstein distance to a vocabulary of up to 50,000 valid words have been used.

Read the paper · More papers on PaperTik