Multi-Modal AI for Structured Data Extraction from Documents
Kiran Kumar Pappula, Guru Pramod Rusum · International Journal of Emerging Research in Engineering and Technology · 2023
Structured data extraction of unstructured documents like scanned pictures, PDF documents, or photos has become a crucial task to accomplish in a wide range of industries in a world that is becoming more and more digitalized. In the following paper, we present a multi-modal artificial intelligence system combining the visual layout analysis with the capability of natural language processing (NLP) to extract structured fields of heterogeneous documents. The offered solution would use convolutional neural networks (CNNs) and transformer-based models to group the interpretation of the spatial layouts, textual contexts, and semantics in a combined manner. The system has proved to be resistant to document formatting inconsistencies, noise, skew, and complex typography by integrating these features. The hybrid architecture initially carries out visual parsing and identifies regions of interest and yields hierarchical layout features. Such features are combined with semantic embeddings trained on pre-trained NLP models like BERT or LayoutLM, allowing the context-aware extraction of fields. The model is trained and tested on the various types of documents in three domains, including insurance claims, billing statements and legal contracts. The performance metrics depict a considerable increase in punctuality and recollected accuracy compared to conventional OCR-based guideline schemes and multimodal one-dimensional models. This study shows the impact of cross-modal reasoning style to resolve the typical obstacles of lacking labels, ambiguous fields, and varying arrangements. The modular structure of the system is also domain-adaptable and extensible, which paves the way for scalable and automated document understanding in enterprise solutions