Enhancing Table Recognition Using Vision Language Models (VLM)

Yunan Liu, Yi Zhu · 2025

Table recognition plays a critical role in transforming unstructured data into structured formats for analysis and decision-making. Traditional methods often struggle with complex table layouts due to challenges such as merged cells, irregular structures, and varying document formats. Recent advancements in large language models (LLMs), particularly Vision-Language Models (VLMs), offer new opportunities for improving table extraction by leveraging both textual and visual information. This study investigates the effectiveness of VLMs in table recognition by evaluating three different approaches. The first approach applies a baseline VLM for table extraction task. The second approach takes a step further by introducing prompt chaining to decompose the extraction process into sequential tasks. The third approach further improves performance by incorporating a table detection step using table transformer to identify table regions before being processed by the VLM. Experimental results show that combining prompt chaining and table detection leads to significant improvements in table extraction accuracy, particularly for complex layouts. The study highlights the potential of VLMs in table recognition applications and automated document processing.

Read the paper · More papers on PaperTik