Autonomous Deblurring Images and Information Extraction from Documents Using CycleGAN and Mask RCNN

Oishee Bintey Hoque, Maisha Binte Rashid, K M Tawsik Jawad · 2020

In this era of automation technology, there has been a resurgence of interest for extracting data from documents so that information can be used more efficiently. The age of storing sensitive information on paper is on the brink of extinction. Almost every organization in the world is shifting towards an efficient cloud storage system. With the global impact of Covid-19 and the norm of “work from home”, the importance of proper data extraction from digital documents has increased manifold. So, in this paper, we proposed a method to extract data from scanned document images by identifying handwritten texts and respective label fields. While scanning images various issues can appear such as - background noises, blurred images due to camera motion, out of focus images, watermarks, stains, or anything that can cause hindrance to readability of users. So, for our method to work better our first approach was to filer the scanned images by removing noises. We trained our dataset on Cycle Generative Adversarial Networks (Cycle-GAN) to generate clean scanned images from noisy images. Later on, for detecting labels and text fields we trained Mask R-CNN on our cleaned dataset. Finally, we extract the information using Tesseract from the detected fields and assign the labels filed with corresponding information on a text file.

Read the paper · More papers on PaperTik