A Word Spotting Method for Arabic Manuscripts Based on Speeded Up Robust Features Technique
Noureddine El Makhfi · Advances in Science Technology and Engineering Systems Journal · 2019
The diversity of manuscripts according to their contents, forms, organizations and presentations provides a data-rich structures.The aim is to disseminate this cultural heritage in the images format to the general public via digital libraries.However, handwriting is an obstacle to text recognition algorithms in images, especially cursive writing of Arabic calligraphy.Most current search engines used by digital libraries are based on metadata and structured data manually transcribed in Ascii format.In this article, we propose an original method of pattern recognition for searching the content of Arabic handwritten documents based on the Word Spotting technique.Our method is both effective and simple, it consists in extracting a set of features from the words we segment in the target images and comparing them with the features of the words in the requested images.The principle of the method is to characterize each word with the Speeded Up Robust Features algorithm whose goal is to find all occurrences of query words in the target image even in the case of low-resolution images.We tested our method on hundreds of pages of Arabic manuscripts from the National Library of Rabat and the Digital Library of Leipzig University.The results obtained are encouraging compared to other methods based on the same Word Spotting technique.