One Hot Encoding and Hashing_Trick Transformation - Performance Comparision
Agata Kozina, Michał Nadolny, Marcin Hernes, Ewa Walaszczyk, Artur Rot · 2024
The purpose of this paper is to compare the effectiveness of two data transformation methods for categorical text variables: one hot encoding and hashing_trick. The analysis was performed using a dataset sourced from leasing companies. This dataset contained a variety of variable types, including textual categorizers, the transformation of which was crucial to the correct operation of the deep learning model. The results of the study are based on a pre-designed experiment consisting of the following stages: transformation by one hot encoding and hashing_trick methods, developing the deep learning model, and statistical analysis of the obtained results. Analysis of the results showed that the two data transformation methods generate statistically different results. The hashing_trick method is distinguished by the higher performance measures of deep learning models. In addition, the tuning of model parameters had a significant impact on the final results, with both data transformation methods producing similar results for some parameter combinations.