Historical Kannada Handwritten Palm Leaf (HKHPL): A Benchmark Dataset for Segmentation of HKHPL Manuscripts

Parashuram Bannigidad, S. P. Sajjan · Cureus Journal of Computer Science. · 2025

Here, we introduce the Historical Kannada Handwritten Palm Leaf (HKHPL) Dataset, the pioneering collection of historical palm leaf manuscripts written in Kannada. It includes original ground-truth and binarised images. This dataset is constructed from randomly selected palm leaf manuscripts collected from various parts of Karnataka and Maharashtra, India, and is publicly available for scientific purposes. The HKHPL dataset significantly contributes to the study of ancient documents and the recognition of handwritten Kannada texts on palm leaf manuscripts, aiding important tasks such as reading characters, segmenting text lines, and analyzing layouts. This dataset addresses challenges posed by ancient manuscripts, including degradation, non-uniform writing surfaces, and handwriting variations. By providing binarised ground truth images, word-annotated images, and isolated character-annotated images, the dataset offers researchers a valuable resource for developing and testing algorithms in image preprocessing, word segmentation, and character recognition. The diverse representations of historical Kannada texts ensure robust machine learning models capable of handling various writing styles, document conditions, and content types. With the HKHPL dataset publicly available, along with its implementation code on GitHub, the creators have opened new avenues for research in digital humanities, historical linguistics, and AI applications in cultural heritage preservation. This dataset can accelerate advancements in the automated transcription and analysis of historical Kannada texts, contributing to the preservation and study of India’s literary heritage.

Read the paper · More papers on PaperTik