Integrating Textual Semantics with Visual Positioning: A Multi-Modal Approach to Enhanced Complex Table Understanding
Yujia Rong, Bo Jiang, Hong Xu · 2025
In the field of project fund supervision, an important research topic is the intelligent data collection and analysis of contract information and financial statements. However, due to the high domain specificity and concise terminology usage in such documents, combined with the frequent occurrence of complexly styled tables, these factors present significant challenges for model-based data processing. This study proposes a multimodal feature fusion approach that integrates textual semantic features with visual positional features from document tables. Our model undergoes specialized pre-training on text-image alignment tasks, enabling simultaneous comprehension of tabular content while establishing explicit associations between textual elements and their spatial positions. The proposed visual-textual semantic feature fusion method has been validated on the publicly available FUNSD dataset. Experimental results demonstrate the effectiveness of our approach in addressing complex table document understanding tasks.