Numerical Embedding of Categorical Features in Tabular Data: A Survey

Qianxi Qiu, Han Liu · 2023

Tabular data, which are also known as structured data, is one of the most common forms of data and can be frequently seen in fields such as healthcare, finance, recommendation systems and network security. In contrast to homogeneous data like image, audio and text, tabular data is usually heterogeneous and shows a small sample size and weak correlations between features. Besides, tabular data may contain not only numerical features but also categorical ones that can not be handled directly by the majority of those popular learning algorithms. The aforementioned characteristics make decision tree ensemble methods, such as random forests and gradient boosting trees, more suitable for processing tabular data than other models. In this context, it is essential to transform categorical features into numerical ones in the setting of category embedding. This paper surveys embedding methods for handling low-cardinality categorical features in tabular data. Specifically, these methods are put into two categories, namely, supervised and unsupervised encoding, which each is introduced in terms of the strategies of category encoding. Moreover, the characteristics of each method are summarized. On the basis of the summarization, further directions of category embedding are suggested within the setting of representation learning.

Read the paper · More papers on PaperTik