Scalable Tabular Hierarchical Metadata Classification in Heterogeneous Structured Large-Scale Datasets Using Contrastive Learning

Bhimesh Kandibedala, Gyanendra Shrestha, Anna L. Pyayt, Todor Ivanov, Michael N. Gubanov · 2025

Tabular metadata (i.e., attributes in a table) identification and classification is a fundamental problem in large-scale data management of structured corpora, especially for complex tables rich in multi-level hierarchical metadata with nesting. Medical, security, data science research literature, Web tables, contain thousands of such complex tables, but often lack or incorrectly label their complex metadata. In this work, we describe an unsupervised, scalable, contrastive-learning approach for classification of multi-layer, hierarchical metadata in such tables. We compared it to the state of the art (SOTA) as well as the latest Large Language Models (LLMs), such as OpenAI GPT 3.5 and 4 with and without Retrieval Augmented Generation (RAG) on several large-scale heterogeneous datasets. We outperform SOTA and LLMs in classifying horizontal metadata (HMD) of deep levels (3–5) and for all levels (1–3) of vertical metadata (VMD). For HMD levels 1–2, SOTA outperforms us insignificantly, with a delta of ≈1%. LLMs with/without RAG slightly outperform us with deltas of 4–5% in accuracy for HMD level 1, but we significantly outperformed LLMs/LLMs+RAG with delta up to 29% for all other levels 2–5 HMD and up to 87% delta for VMD.

Read the paper · More papers on PaperTik