Interaction Modeling for High-Dimensional Mixed Data with Numerical and Categorical Features

Sara Hoshino, Taiki Ito, Takuya Matsuzawa, Haruki Ozawa, Sumika Arima · Procedia Computer Science · 2024

This study aimed to develop an interaction modeling method for high-dimensional industrial data with sparsity. Particularly, we discussed the potential and limitations of Sparse Factorization Machines (SFM) with feature selection capabilities after examining the applicability of Factorization Machines (FMs) to numerical and categorical mixed data like industrial data. FMs has been already a major recommendation engine outperforms SVM because of robustness to the sparse data. However, conventional FMs and SFM based on the L2 norm regularization tolerate huge False-Positives (FPs), which is fatal in the application to real data to which the oracle model is unknown. Therefore, in this study, we focus on the way to automatically reduce the several millions of FPs in interactions while keeping high True-Positives (TPs). For the purpose, SFM with trigonometric inequality (TI) upper boundaries (Atarashi et al., 2021) is improved by two directions. The first is the development of TI_SFM (L1) with an L1 norm for selection of main factors, particularly for FPs reduction of the main factors. The second is the application of adaptive technique for reducing FPs of interactions (combinatorial features). We newly developed “Adaptive SFM” with adaptive technique to introduce data-driven penalty of the interaction term. As the result of numerical evaluations using a mass production oracle interaction model and several simulation data, False-Positives of the main factors ( F P main ) and the interactions ( F P int ) are significantly reduced, while keeping high level of True-Positives of the main factor (TP main ) and the interactions (TP int ). Concretely, our proposed Adaptive SFM (L1) outperforms the original TI_SFM (L2) as much reducing F P main and F P int by over 99% when applying our proposed penalty considering not only the relationship between the explanatory variable X i and the objective variable Y, as basic adaptive technique, but also the factor loading indicating the relationship between the latent vector internally optimized and the explanatory variable X i . Our contribution is to lead the possibility for applications to actual industrial data with uncertain main factors and interactions, beyond the applications as recommendation engines.

Read the paper · More papers on PaperTik