Zero-Shot Invoice Information Extraction Using Foundation Models with Spatial Prompt Tuning
Ranadheer Reddy Charabuddi · International Journal of Intelligent Systems and Applications in Engineering · 2025
Extracting structured information from scanned invoices poses significant challenges due to diverse layouts, linguistic variability, and the scarcity of annotated training data. To address this, the study introduces a zero-shot invoice information extraction framework that leverages the Donut foundation model, integrated with spatial prompt tuning. Unlike conventional OCR-based pipelines, the proposed approach operates directly on document images without the need for explicit text recognition or task-specific fine-tuning. The method was evaluated using the SROIE v2 dataset, comprising 973 annotated invoice images, and was implemented using the Python framework. Spatially-aware natural language prompts were used to guide the model’s attention toward relevant regions such as headers or totals. Experimental evaluation demonstrated a notable performance gain, with the model achieving 98.0% accuracy, surpassing baseline methods like BiLSTM-CRF and LayoutLM by over 4%. The results validate the model’s effectiveness and scalability for real-world document automation, especially in zero-shot settings with high template variability. DOI: https://doi.org/10.17762/ijisae.v13i1s.7722