Systematic construction of anomaly detection benchmarks from real data

Andrew Emmott, Shubhomoy Das, Thomas G. Dietterich, Alan Fern, Weng‐Keen Wong · 2013

Research in anomaly detection suffers from a lack of realistic and publicly-available problem sets. This paper discusses what properties such problem sets should possess. It then introduces a methodology for transforming existing classification data sets into ground-truthed benchmark data sets for anomaly detection. The methodology produces data sets that vary along three important dimensions: (a) point difficulty, (b) relative frequency of anomalies, and (c) clusteredness. We apply our generated datasets to benchmark several popular anomaly detection algorithms under a range of different conditions.

Read the paper · More papers on PaperTik