An Approach to Automatically Extract Predictive Properties from Nominal Attributes in Relational Databases

Valentin Kassarnig, Franz Wotawa · 2018

Feature engineering is a fundamental step in data mining and yet it is both difficult and expensive. Hand-crafting features is not only a time-consuming task that requires specific domain knowledge, it also may prevent new information to emerge. The extraction of meaningful features from relational data is particularly difficult due to complex relationships between tables. In the last decade there is an emerging trend towards automating the process of constructing propositional features from relational data and such approaches have been successfully used for solving numerous real-world problems. Despite their success, most of them lack an adequate support of nominal attributes. We present a new approach helping propositionalization methods to extract meaningful features from nominal attributes and improve their predictive performance. In an experimental evaluation on three datasets we demonstrate that the proposed technique is capable of producing novel features that are highly correlated with the target attribute. Furthermore, those features can reveal relationships among the distinct categorical values allowing to compare and order them. Finally, experimental results show that those new features can significantly improve the predictive performance in classification tasks.

Read the paper · More papers on PaperTik