Word Extraction from Arabic Handwritten Documents Based on Statistical Measures

Ayman Al-Dmour, Raed Abu Zitar · International Review on Computers and Software (IRECOS) · 2016

In Arabic, word extraction is particularly challenging because words are often divided into sub-words, and a few letters do not connect to the following letter. In this paper, we present an efficient method for extracting words from Arabic handwritten documents. The proposed method is based on two groups of spatial measures (the lengths of connected components (CCs) and the gaps between these CCs) which differentiate successive CCs in text lines. Lengths are clustered into three distinct clusters to identify an optimal threshold for separating isolated letters, sub-words, and words. Besides, Gaps are clustered into two clusters, to indicate whether the gap occurs "between-words" or "within-a word". This clustering is implemented using Self-Organizing Map (SOM) algorithm. The efficiency of the proposed method was tested by conducting experiments on 35 ages of handwritten Arabic text, accessed from benchmarking Database for Arabic Handwritten Text Recognition Research (AHDB). Our tests produced very promising results, achieving a correct extraction rate of 86.3%.

Read the paper · More papers on PaperTik