Character and Word Level Recognition from Ancient Manuscripts using Tesseract

Akshay Menon, Rohit Sreekumar, B J Bipin Nair · 2023

In uses adapti ve Gaussian thresholding and pre -processing techniques like binarization to improve the quality of the images. Additionally, segmentation is carried out at the line-and word-levels utilizing the blob approach and regional zoning. Recognizing the characters and creating the database are the last steps. The efficiency of the suggested strategy for successfully identifying and digitizing Malayalam manuscripts is demonstrated by the achievement of recognition accuracy of above 93%. The goal of this project is to contribute to the creation of a comprehensive Malayalam database that will aid in the advancement of natural language processing research and development.

Read the paper · More papers on PaperTik