Word Extraction from Arabic Handwritten Documents Based on Statistical Measures
Ayman Al-Dmour, Raed Abu Zitar · International Review on Computers and Software (IRECOS) · 2016
In Arabic, word extraction is particularly challenging because words are often divided into sub-words, and a few letters do not connect to the following letter. In this paper, we present an efficient method for extracting words from Arabic handwritten documents. The proposed method is based on two groups of spatial measures (the lengths of connected components (CCs) and the gaps between these CCs) which differentiate successive CCs in text lines. Lengths are clustered into three distinct clusters to identify an optimal threshold for separating isolated letters, sub-words, and words. Besides, Gaps are clustered into two clusters, to indicate whether the gap occurs "between-words" or "within-a word". This clustering is implemented using Self-Organizing Map (SOM) algorithm. The efficiency of the proposed method was tested by conducting experiments on 35 ages of handwritten Arabic text, accessed from benchmarking Database for Arabic Handwritten Text Recognition Research (AHDB). Our tests produced very promising results, achieving a correct extraction rate of 86.3%.