Towards Accurate Recognition of Historical Arabic Manuscripts: A Novel Dataset and a Generalizable Pipeline

Hakim Bouchal, Ahror Belaid, Farid Meziane · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

In today’s digital world, we are committed to digitizing thousands of handwritten transcriptions to preserve their content. Historical Arabic Handwritten Text Recognition (HAHTR) remains a challenge for computer vision systems, due to the many difficulties inherently associated with document image quality and the complexity of Arabic script. In this work, we address the problem of recognizing historical Arabic documents that adapts to different writing styles and degrees of legibility. We developed a system that is able to recognize a whole page of a historical Arabic handwritten text in two consecutive steps comprising text line detection and recognition he proposed approach performs detection using bounding boxes followed by a neural network-based model for character-level text recognition. However, the lack of data hinders the mass digitization of Arabic historical documents. Therefore, we provide a new and freely available dataset, focusing on diverse handwriting styles to facilitate a strong generalization of the trained model. This dataset will significantly benefits researchers and practitioners by accelerating progress in the field of HAHTR. Extensive experimental work demonstrates that the recognition models are effective when trained with different sources of data, and having different writing styles does not penalize the model’s ability to generalize but rather enhances it. Additionally, we define and develop a new metric to evaluate model robustness against character misclassification, particularly for characters with similar patterns. The experiments conducted demonstrated that the proposed HAHTR pipeline is accurate and highly generalizable, as well as the validity of bounding box methods for detecting text lines. The training approach with different data sources enabled us to surpass the state-of-the-art results with 5.7% of Character Error Rate (CER) on the KHATT database.

Read the paper · More papers on PaperTik