A Data Entry Optical Character Recognition Tool using Convolutional Neural Networks
Samarth Ghulyani, Dhyanendra Jain, Prashant Singh, Sarthak Joshi, Ankit Ahlawat · 2022 IEEE IAS Global Conference on Emerging Technologies (GlobConET) · 2022
Almost all institutions and organizations rely substantially on data to run their operations. Data is necessary for making informed decisions, adapting to change, and defining strategic objectives. Data administration has always relied on manual data entry. Manual input is used to entail transferring data from various documents into record books, ledger books, and other such books. Manual data entry, as used in recent years, comprises manually entering particular and predetermined data into a target program, such as customer name, business kind, money amount, and so on, from various sources, such as paper bills, invoices, orders, receipts, and so on. Depending on the sort of business, the target program can be handwritten records, spreadsheets, or computer databases. Several businesses require manual data entry, which has a high rate of mistakes. This is because the manual approach places far too much reliance on the ability of humans tocomprehend handwritten documents. As a result, a method for retrieving and storing information from images, particularly text, is required. OCR (optical character recognition) is a rapidly growing topic of research aimed at creatinga computer system that can automatically extract and interpret text from images. OCR converts any type of text or text-containing documents, such as handwritten text, printed text, or scanned text images, into an editable digital format for deeper and more complex processing. As a result, OCR enables a machine to recognize text in such documents without the need for human intervention. In order to achieve successful automation, a few significant difficulties must be identified and resolved. One of the most pressing issues is the quality of character typefaces in paper documents, as well as image quality. The computer system may not correctly recognize characters as a result of these difficulties. We examine OCR utilizing four different approaches in this research. We begin by laying down all of the possible issues that may arise during the OCR stages. We next go over the pre-processing, segmentation, normalization, feature extraction, classification, and post- processing aspects of an OCR system. As a result, this conversation paints a rather complete picture of the current state of text recognition domain.