Optical Character Recognition using Tesseract Engine

Nikita Kotwal, Gauri Unnithan, Ashlesh Sheth, Nehal Kadaganchi · Zenodo (CERN European Organization for Nuclear Research) · 2021

The technology of optical character recognition (OCR) was used to transform printed text into editable text. In a variety of applications, OCR is a very helpful and popular approach. Text preparation and segmentation techniques can influence OCR accuracy. Because of the image's varying size, style, orientation, and intricate backdrop, retrieving text from it might be challenging at times. We begin by discussing the Optical Character Recognition (OCR) technology, its design, and the experimental results of OCR conducted by Tesseract on medical data images. We end this work with a comparison of this tool with other detection methods in order to improve detection accuracy.

Read the paper · More papers on PaperTik