Document Analysis Using Adaptive Hybrid Deep Learning Techniques
Koushik Sundar, M. Bhavani, M Jaeyalakshmi, R Vijayakumar · 2024
A new era of massive and varied data repositories has been brought about by the digitization of documents, necessitating the creation of reliable systems for interpreting and classifying scanned documents. Using deep learning capabilities, this study investigates state-of-the-art methods for document analysis and classification. Optical character recognition (OCR) systems have been revolutionized by deep learning models, especially CNN and transformer-based architectures. These algorithms are remarkably accurate at extracting text from scanned documents, independent of language, font size, or style. Transformer models perform well in comprehending textual material, whereas CNNs are skilled at extracting discriminative features from document images. Utilizing transfer learning approaches to maximize model performance while minimizing the requirement for large amounts of labelled data has been the subject of recent study. Deep learning has also resulted in notable advancements in Named Entity Recognition within scanned documents. Document organization and comprehension have been made easier by the remarkable accuracy with which names, dates, and locations can be extracted by bi-directional LSTMs, BERT-based models, and their variations. Improved document processing, organization, and accessibility are made possible by the approaches covered here, which have a wide range of applications in the financial, medical, and legal services sectors.