Optical character recognition: transforming images into text

Nikhil Kushwaha, Om Asati, Mainak Sadhya · 2024

Optical character recognition (OCR) is a technology that enables the full recognition of the alphanumeric printed characters. It scans the text image then interprets the printed characters and converts it into a corresponding editable text document. This chapter describes the working of the OCR to extract the text content of the image and the implementation of the convolutional neural network to extract the features of the image. ResNet (Residual Network) is a deep learning architecture that addresses the problem of degradation and vanishing gradients in very deep neural networks. Optical character recognition can also be achieved using different programming languages like Java, C++, C#, JavaScript, Ruby, MATLAB, etc., but because of less complexity, less syntax and easy to use functionalities of a bunch of diverse libraries, Python is the most used language in such fields as image processing and video processing.

Read the paper · More papers on PaperTik