Optical Character Recognition Systems for Accurate Interpretation of Handwritten Telugu Scripts

R Prasanna Kumar, D. Spoorthi Keya, M Charan Kumar · 2024

Visual data dominates the digital landscape these days. However, there remains a significant gap in the digitization of textual content in regional languages such as Telugu. This research addresses this gap and utilizes the capabilities of the programming language and set of specialized libraries pytesseract for optical character recognition (OCR), PIL for image processing and Levenshtein for accuracy evaluation to appropriately extract text from images in the Telugu script. There are several important steps in this project. Initially, this work sets up a dependable system for uploading pictures. Secondly, carefully extracting text from these submitted photos using py-tesseract. Using the Levenshtein distance measure, and then compared with the ground truth to determine correctness. This metric measures the accuracy of the text output through looking at the slight variations between the word error rate (WER) and character error rate (CER). This is capable of finding misspelled words and unknown characters providing a precise evaluation of the OCR’s effectiveness. This study also examines the current state of OCR in Telugu text processing and indicates the benefits of current technologies for locating and converting Telugu text discovered in image documents. It also demonstrates improvements in OCR technology.

Read the paper · More papers on PaperTik