DB-Outlier Detection by Example in High Dimensional Datasets

Yuan Li, Hiroyuki Kitagawa · 2007

Outlier detection is an important problem with applications in many fields. Such applications generally process high dimensional datasets. Among the existing methods of detecting outliers, Distance-Based outlier (DB-Outlier) detection is one of the most commonly used and simplest approaches, since it detects outliers only by calculating distances between data points. However, in high dimensional space, data is sparse, so every data point becomes a good outlier candidate. A Subspace-Based method has been proposed to deal with the curse of dimensions. It shows that meaningful outliers are likely to be identified by examining the behavior of data in low dimensional projections. On the other hand, most existing methods detect outliers with parameters being determined by users in advance. Such parameters usually contain hidden user view of outliers. Example-Based outlier detection methods are presented to be promising in discovering the hidden user view of outliers. In this paper, we discuss a new technique to detect DB-Outliers in high dimensional datasets based on user examples. Our proposed method makes use of Subspace-Based and Example-Based methods to discover a subspace where user examples are outstanding more significantly than in any other subspaces, and reports DB-Outliers detected in this subspace.

Read the paper · More papers on PaperTik