VTLL: Visual-Texture Multi-Modal Fusion for Table Structure Recognition Based on Logic Location Regression

Jun Fang, Chongyang Zhang · 2024

Tables play a crucial role in both documents and daily life, thus sparking significant interest in the research of automatic table structure recognition(TSR). Recent methods primarily achieve recognition by predicting makeup sequences or adjacency relationships between table cells. However, these methods have some limitations in applications. The former requires additional computation of heuristic rules to recover the structure of the table, which not only increases the computational complexity but also may affect the recognition accuracy. The later relies on huge training data and inefficiency decoders, which not only raises the cost but also limits their application in real-time or large-scale data processing scenarios. At the same time, they often ignore the significance of cell logical location and fail to effectively utilize the rich text information that naturally exists in the table. In this paper, we propose a new framework called VTLL to solve the problem. The method extracts visual and textual features to do adaptive feature fusion, and we also introduce a cascading regressors to predict the fused features multiple times, while combining intra-cell and inter-cell losses. As for the regression header part, we use a mixed context aggregator to understand the inter-cell relationships. We evaluated the proposed methods on several public datasets, experiment results demonstrate that VTLL performs better.

Read the paper · More papers on PaperTik