Info-ExtNet: An Information Extraction framework from Noisy Structured Documents

Sk Shamim Aktar, Debapriya Banik · 2025

The exponential growth of digital data necessitates efficient and accurate information extraction (IE) from structured documents. Traditional rule-based systems, limited by domain-specific constraints, struggle with scalability and adaptability. This paper proposes Info-ExtNet, an innovative model that leverages the LayoutLM architecture by integrating text, 2D positional, and image embeddings for enhanced document understanding. Our approach operates through three key stages: preprocessing, modeling, and post-processing. Advanced Optical Character Recognition (OCR) techniques and layout parsing are used during preprocessing to transform scanned documents into structured data while retaining essential spatial information. The core modeling phase employs a transformer-based architecture with multimodal embeddings to capture both textual and visual features effectively, and the post-processing stage refines the extracted data to ensure precision and consistency.Our model was evaluated using benchmark datasets such as FUNSD and SROIE, achieving high performance metrics, including a Precision of 0.94, Recall of 0.94, and an F1-Score of 0.95. These results surpass state-of-the-art (SOTA) methods, demonstrating the robustness of our model in handling complex document layouts. The extracted information, saved in a structured CSV format, supports immediate integration into various downstream applications. This work provides a scalable solution for document intelligence, paving the way for future research in document processing and analysis.

Read the paper · More papers on PaperTik