Complex Encoding
Kassymzhomart Kunanbayev, Islambek Temirbek, Amin Zollanvari · 2021
A salient problem in machine learning is transforming categorical variables into efficient numerical features. This focus is warranted due to the ubiquity of categorical data in real-world applications but, on the contrary, the development of many machine learning methods based on the assumption of having numerical variables. Perhaps the most popular existing categorical to numerical conversion techniques include one-hot encoding, thermometer encoding, and ordinal (integer) encoding. One-hot and thermometer encodings become computationally inefficient and may lead to ill-conditioning for high-cardinality categorical variables as they create high-dimensional data matrices. As ordinal encoding does not change data dimensionality, it is significantly more memory-efficient; however, the spurious ordering induced by ordinal encoding among unordered categorical values can hamstring the performance of constructed predictive models. In this paper, we propose a new encoding technique, namely complex encoding, that provides a symmetric representation of categorical values in the complex plane. To show the efficacy of the complex encoding in terms of error rate and memory usage, we conducted a set of numerical experiments using real datasets and the family of linear discriminant functions for complex Gaussian distributions. Empirical results show that not only complex encoding avoids the ill-conditioning problem of one-hot and thermometer encodings, it can generally lead to a comparable or higher classification accuracy with respect to others (when they are applicable) at the expense of only about two-fold increase in memory usage with respect to ordinal encoding.