Interaction Modeling for High-Dimensional Mixed Data Considering Distribution and Correlation
Taiki Ito, Takuya Matsuzawa, Sumika Arima · Procedia Computer Science · 2025
This study aimed to develop an interaction modeling method for high-dimensional mixed data with numerical and categorical variables, assuming application to industrial data. In particular, we focused on the development of a method for the selection of interactions consisting of numerical variables and categorical variables, which have not yet been discussed in previous research. Specifically, we proposed an advanced integration form of Adaptive Sparse Factorization Machines (SFM1A) and Safe Pruning, and confirmed the effect through verification of data simulating real industrial data. As a background, the conventional Factorization Machines based system allows incorrect selection due to a manner of recommendation systems, and tens of thousands of incorrect interactions remain even in SFM1A, which reduces the number of incorrect interactions by 97%, so it has been a fatal limit in the application phase to real data where the correct variables are unknown. Therefore, in this study, we proposed SFM1A_SPC in which the SPC criterion of Safe Pruning, which is an effective screening method for interactions between categorical variables, is applied as a pretreatment to SFM1A. As a result, the selection of incorrect interactions was reduced by 99.7% for data with only categorical variables, and the practical application target (a few hundred or less) was expected to be achieved. On the other hand, the method of converting numerical variables to categorical variables by the simple approach of 0-1 normalization described in the original SPC paper for mixed data with numerical and categorical variables could not be reduced to the target level for practical use. Therefore, as an approach to mixed data with numerical and categorical variables, we propose SFM1A_NCmix in which numerical variables are appropriately categorized and normalized before SPC criterion are applied with data-driven reasonable relaxation. Numerical evaluations are performed on mixed data with numerical and categorical variables of any three types that consider correlations and distributions. As a result, the F1 score was improved by several times (0.90 or more) when using the SFM1A_NCmix for all datasets. In particular, for the mixed data with numerical and categorical variables that follow a multivariate normal distribution that consider correlations, it achieves the incorrect interactions ( FP int ) level for practical applications, besides the superiority over SFM1A and SFM1A_SPC.