A dataset of Warped Historical Arabic Documents

Ali Dulla · 2018

Historical documents are considered one of the most important human wealth and a source of intellectual production. Unfortunately, due to aging effects, multiple noises and arbitrary geometric distortions are found in the document image. This paper presents a novel dataset (and the methodology used to create it) based on an extensive variety of historical Arabic documents containing clean information basic and homogeneous-page layouts. The tests are executed on printed and handwritten documents acquired respectively from some imperative libraries, for example, Qatar Digital Library, the British Library and the Library of Congress. We have collected and commented on 200 archival document images from various sites and time periods. It is based on different documents from the 16th19th century. The dataset involves varying page layouts and degradations that challenge text line segmentation techniques. Ground truth is created using the Aletheia tool by PRImA and stored in an XML representation, in the PAGE format. The dataset offered will be effectively accessible to specialists worldwide for inquire about into the impediments confronting different historical Arabic documents, for example, geometric correction of historical Arabic documents.

Read the paper · More papers on PaperTik