On the influence of categorical features in ranking anomalies using mixed data
Mathieu Garchery, Michael Granitzer · Procedia Computer Science · 2018
Most unsupervised anomaly ranking approaches are compatible with numeric data only, leading to categorical features often being ignored in practice. Even though some methods address this issue, few support mixed data and the influence of excluding or including categorical attributes has not been studied well yet. In this paper, we take a first step towards considering categorical and numeric attributes jointly for unsupervised anomaly ranking by benchmarking selected methods. We introduce three new approaches: two entropy-based methods based on individual and collective entropy contribution, as well as an extension of Isolation Forest supporting mixed data, and benchmark them against SPAD, a state-of-the-art probabilistic anomaly ranker. We observe that our entropy methods detect very similar anomalies in practice, and these anomalies are mostly globally isolated observations. Both entropy methods are also closely related to SPAD. Our empirical study additionally shows that categorical features can have high impact on anomaly ranking performance and thus should not be blindly ignored.