Evaluating Transformers for Tabular Data: The Impact of Tokenization, Dataset Characteristics and Noise

Goran Oreški · 2025

Transformers have achieved state-of-the-art performance in natural language processing and computer vision; however, their applicability to tabular data classification remains an open research question. One of the primary challenges arises from the tokenization process, which disrupts the inherent relationships between numerical and categorical attributes. This study proposes a new positional encoding scheme and systematically evaluates the impact of dataset characteristics, including: size, feature composition, noise, and redundancy, on the performance of transformer-based models. A synthetic data set generator is used to create controlled environments, allowing for precise variations in dataset characteristics.Key findings indicate that proposed feature-based positional encoding significantly improves learning efficiency. Additionally, transformers architecture exhibit a degree of robustness to noise, as models trained on noisy datasets eventually recover performance after extended training. However, an excessive number of attributes, particularly high-cardinality categorical features, hinders learning, necessitating feature aggregation or dimensionality reduction.The study highlights the conditions under which transformers could be a viable alternative to traditional machine learning models for tabular data classification.

Read the paper · More papers on PaperTik