Encoding Techniques for High-Cardinality Features and Ensemble Learners

Justin M. Johnson, Taghi M. Khoshgoftaar · 2021

This study evaluates the classification performance of five encoding techniques for high-cardinality categorical features. Encoding techniques are tested using popular bagging and boosting ensemble methods on the latest Medicare Part B fraud classification data set, where the healthcare procedure code feature includes 7,752 unique values. To provide an additional baseline, we also evaluate these encodings on the multilayer perceptron. One-hot encodings are compared to a baseline aggregated encoding that excludes the procedure code feature to determine if the categorical feature significantly affects performance. Next, LightGBM and CatBoost's built-in strategies for categorical feature handling are compared to Hcpcs2Vec embeddings, distributed representations of procedures that encode semantic similarities. Statistical tests show that the inclusion of the categorical feature significantly improves performance for all ensemble learners when a one-hot representation is not used. Results also show that the XGBoost learner with Hcpcs2Vec encodings perform best overall with an average AUC of 0.8715. Our comparison of diverse encoding techniques for the high-dimensional categorical feature makes this study a unique contribution in the areas of ensemble learning and healthcare fraud prediction.

Read the paper · More papers on PaperTik