Fine-Grained Layout-Aware Information Extraction for Imaged Documents
Yiliang Chen, Yaming Yang, Zhiling Zeng, Hang-Bom Bu, Chentai Yun, Bo Wen Wei, Kangjun Liu, Chuheng Liang, Dongming Lin, Maolin Xu · 2024
Visual information extraction is becoming more and more important in the field of document artificial intelligence, such as recognizing various types of semantic text entities and extracting inter-entity relationships in documents such as business contracts, medical reports, etc., which reflects its key role in processing complex document information. Currently, there are mainly Transformer-based and graph neural network-based approaches to realize semantic entity recognition and relationship extraction. Among them, Transformer-based methods are more commonly used and have achieved certain results, but they rely on large-scale labeled datasets and have high training costs. The graph neural network-based method can selectively associate information, with the significant advantages of lightweight and high efficiency, but fails to effectively utilize the key information contained in the form layout in the process of node selection and edge connection. In this paper, inspired by the related work on graph neural networks, we propose a fine-grained layout-aware image-based document information extraction method. First, finegrained document graph nodes are constructed to accurately obtain cell information and multi-dimensional feature encoding. Second, the redundant connections are effectively eliminated with the help of document layout logic, and the layout-aware node association is realized. Finally, many experiments are carried out on FUNSD and XFUND Chinese datasets, and the experimental results show that the proposed method exceeds the existing information extraction methods and achieves 62 % of the F1 value, which provides a strong support and a new direction for the research of document information extraction.