Text Information Extraction from Digital Image Documents Using Optical Character Recognition

Md. Mijanur Rahman, Mahnuma Rahman Rinty · 2023

In a digital world, different images are generated from various sources and used in multiple applications. Text present in these images contains meaningful evidence for explanation and structuring of images and the semantic understanding of the images. But text extraction is challenging due to variations of text patterns (image types, mode of capture, text position, etc.). Thus, text information extraction from digital documents is an emerging area of research. Various systems have been reviewed for the detection and extraction of text from images. This chapter is concerned with developing optical character recognition based text information extraction from digital image documents. The optical character recognition system involves several algorithms utilized for text detection and identification, text localization, text segmentation, and categorization of extracted features. Tesseract is currently the most accurate optical character recognition engine applied to design the proposed system and implemented using the C/C++ programming paradigm in the Linux platform. The experimental results demonstrated a significant performance for standard document images and scanned pages from books/magazines.

Read the paper · More papers on PaperTik