Elicitation of Candidate Subspaces in High-Dimensional Data
Srikanth Thudumu, Philip Branch, Jiong Jin, Jugdutt Singh · 2019
Anomaly detection is an important research area in data mining and has been studied intensively in recent years. The increasing number of features, non-informative noises and other irrelevant features, makes it challenging to detect anomalies in high-dimensional data. When analyzing high-dimensional data, anomalies are difficult to identify due to the sparsity caused by the curse of dimensionality. One most commonly used algorithm for reducing dimensionality is Principal Component Analysis (PCA), however, it is known to be sensitive in identifying anomalies. Furthermore, anomalies are rare and only can be found in low-dimensional subspaces. To effectively find anomalies in the high-dimensional space, we propose a technique that explores locally relevant and low-dimensional subspaces where anomalies may be possibly hidden due to the sparsity caused by the curse of dimensionality, and we call these subspaces as candidate subspaces for anomalies. In particular, the proposed technique integrates a Pearson Correlation Coefficient (PCC) and PCA, thereby combining the highest variances. Our experimental results showed that the technique gives good results when anomalies are synthetically introduced.