Intelligent Document Finding using Optical Character Recognition and Tagging

AqeelHussein Abbas, M. Syed Shahul Hameed, S. Balakrishnan, Kaliyaperumal Sugirthamani Anandh · 2022 International Conference on Automation, Computing and Renewable Systems (ICACRS) · 2022

In the era of digitalization, the assortment and exploration of great volumes of documents is becoming progressively significant for enterprises to increase their productions and practices. Optical Character Recognition (OCR) is a procedure of identifying text in scanned (image-based) documents. This paper aims to deliver seamless searching of documents in file systems using Optical Character Recognition (OCR) and Natural Language Processing (NLP). Our paper includes the following phases: "Text Identification (in terms of text files), Image Capturing, Image Enhancement, Image Identification, OCR, Data Extraction and Quality Assurance". In case of text files, the data extraction is done in the first phase itself. The document management system "processes both structured document images (ones which have a standard format) and unstructured document images" (ones which do not have a standard format). In the tagging phase, the document is divided into segments and the tags for each segment are generated using Natural Language Processing.

Read the paper · More papers on PaperTik