Benchmarking TabLM: Evaluating the Performance of Language Models Against Traditional Machine Learning in Structured Data Tasks

Wajiha Abdul Shakir · 2024

This study presents a comprehensive benchmarking of TabLM, a language model derived from DistilBERT, against traditional machine learning models such as Support Vector Machines (SVM), Light Gradient Boosting Machine (LGBM), and XGBoost, as well as deep learning models like TabNet and TabPFN. We evaluate performance across datasets including Titanic, Iris, Wine, and Diabetes, considering feature selection, scaling, and missing data imputation. The results show that TabLM achieves an accuracy of 72% on the Titanic dataset, while SVM and LGBM outperform it with accuracies of 78% and 80%, respectively. Precision, recall, and F1-scores were also measured, with SVM achieving a precision of 0.78 and an F1-score of 0.77, compared to TabLM’s precision of 0.72 and F1-score of 0.70. These findings suggest that traditional models continue to excel on datasets with fewer than 1000 samples, while feature selection enhances TabLM’s performance.

Read the paper · More papers on PaperTik