Improving Classification of Scanned Document Images using a Novel Combination of Pre-Processing Techniques
Arpana Dipak Mahajan, Selvakuberan Karuppasamy, Subhashini Lakshminarayanan · 2023
The accuracy of the optical character recognition (OCR) outputs is critical to the success of real-world applications. It is necessary to preprocess an image before an OCR system can correctly interpret the information contained therein. The technique presented employs a novel combination of preprocessing techniques to prepare document images for OCR. Text that has been extracted may be used in a range of ways, including business intelligence operations that often involve the extraction of meaningful semantic information from massive volumes of documents for subsequent strategic decision-making activities. This paper demonstrates a unique combination of image preprocessing on scanned documents which can be leveraged for semantic document understanding using LayoutLMv2. In the RVL-CDIP dataset's advertisement class, this study has included a few more samples. The proposed objective is to show that text classification systems can classify scanned documents accurately after applying proper preprocessing techniques. On the benchmark RVL-CDIP dataset, investigation indicates that the presented methodology outperforms the current state-of-the-art method by 2.34% of classification accuracy, achieving state-of-the-art results for document image classification.