GVDIE: A Zero-Shot Generative Information Extraction Method for Visual Documents Based on Large Language Models
Senmao Qi, Fei Wang, Hongzhi Sun, Yang Ge, Bo Justin Xiao · 2024
Document Information Extraction aims to extract entities and relationships from visually rich documents. Traditional methods require significant annotation and lack generality. In this paper, we propose GVDIE, a Generative Visual Document Information Extraction method that leverages large language models for zero-shot information extraction. This method aligns document-level multimodal features with text layout features, enhancing performance. We introduce a layout augment module for key text block features and transform the document information extraction from sequence labeling task to a generative one. This approach improves flexibility, efficiency, and adaptability in processing diverse, unstructured documents.