Classification of Table Cells Based on LLM Prompts

Mengjie Liu, Chenyang Bu, Shengxing Bai, Bingbing Dong, Xindong Wu · 2024

Tables, as an important means of data storage, are widely used in spreadsheets, web tables, and PDFs. By integrating information from table data with knowledge re-trieved from an external knowledge base, and examining the correspondences between cell values in the table and instances in the knowledge base, we can extract knowledge from the table to augment and enrich the knowledge base. To achieve this goal, we first need to classify table cells based on their functions in the layout. Due to the diverse structures arising from the arrangements of rows and columns, as well as the complexity of content resulting from concise data storage, current automation techniques heavily rely on stylistic features of table cells, such as font or color. Moreover, these methods are rarely experimented with or validated on tables without style features. Recent literature indicates that large language models (LLMs) demonstrate an ability to understand the structure and content of tables in tasks such as table judgment reasoning. Even without extensive feature inputs or pre-training, LLMs still show comparable results to machine learning and deep learning in these tasks. Therefore, this paper attempts to apply LLMs to table cell classification without using other stylistic features. We have designed a 4-component prompt paradigm (Classification Definition, Instruction, Table, Com-pletion), representing respectively the classification definition, task instructions, table data, and result output. We conduct experiments on three datasets CIUS, SAUS, and DEEX for table cell classification with one-shot learning. Our experimental results show that with the assistance of LLMs, better results can be achieved without utilizing stylistic features.

Read the paper · More papers on PaperTik