Official Document Text Extraction using Templates and Optical Character Recognition

Florin Harbuzariu, Cosmin Irimia, Adrian Iftene · 2023

Documents have been used across history ever since civilized societies first began appearing. Documents are used everywhere today in our daily activities and were affected by technological leaps. From documents written on paper, we switched to digital documents. One of the technological fields that are dealing with documents is Computer Vision, specifically OCR, or optical character recognition. OCR is the process in which an image containing text is converted into digital text format [1]. Because computers are used everywhere nowadays, systems have already been designed for working with documents. In many systems that deal with documents, there is still a need for manual work. This paper proposes a way in which OCR can be applied to official documents for the extraction of their text.

Read the paper · More papers on PaperTik