Very Fast Outlier Detection in Large Multidimensional Data Sets.

Amitabh Chaudhary, Alexander S. Szalay, Andrew Moore · 2002

Outliers are objects that do not comply with the general behavior of the data. Applications such as exploration in science databases need fast interactive tools for outlier detection in data sets that have unknown distributions, are large in size, and are in high dimensional space. Existing algorithms for outlier detection are too slow for such applications. We present an algorithm based on an innovative use of k-d trees that doesn't assume any probability model and is linear in the number of objects and in the number of dimensions. We also provide experimental results that show that this is indeed a practical solution to the above problem.

Read the paper · More papers on PaperTik