The HWDI Dataset of Camera Captured Warped Hindi Text Document Images

Vibhanshu Singh Sindhu, Yash Sant, Ruchika Malhotra, S. Indu · 2022 12th International Conference on Cloud Computing, Data Science & Engineering (Confluence) · 2022

Nowadays, capturing text images and using the information contained in these texts have become an integral part of almost everyone’s life in one way or the other. Although a myriad of datasets and advancements have flourished for the Alphabetic scripts such as English for text recognition, text line segmentation, and extraction processes, not much has been developed for Alpha Syllabary scripts such as Hindi. For instance, there is no publicly available dataset of warped text document images for Alpha Syllabary script Hindi for the use of community to work and experiment on. Hence, as a first move in the progression of the above-mentioned field, this paper introduces a novel dataset (HWDI - HINDI WARPED DOCUMENT IMAGES) of camera-acquired text document images in Alpha Syllabary script Hindi and the methodology employed to produce it. The dataset contains 253 warped images of varying warping, having convex or concave surfaces and additionally differs in the area where warping is present, namely, left warping, right warping, and middle warping that challenge the existing text line segmentation techniques. We have also provided ground truths and flatbed-scanned images of the warped images along with the dataset for the use of researchers for various purposes. This innovative dataset will benefit research society to further take ahead the development in the field of text line segmentation, text recognition, dewarping, text extraction and many more developments for Hindi language.

Read the paper · More papers on PaperTik