Universal Anomaly Detection and Applications

Sehong Oh · Deep Blue (University of Michigan) · 2023

Anomaly detection is important in many research areas including fraud detection and biological change detection. However, anomaly detection is a difficult task due to the lack of anomalies available for training. In this thesis, we propose a compression-based nonparametric anomaly detection method for time series and image data using a pattern dictionary. This method constructs two features (typicality and atypicality) to distinguish anomalies based on normal training data captured in a tree-structured data structure. The typicality of a test sequence is a measure of how well the data can be compressed by the pattern dictionary. The typicality can be used as an anomaly score to detect anomalous data at a certain threshold. The atypicality of a sequence is a measure of compressibility of the test data by a universal source coder, determined independently of training data. The typicality and the atypicality of each sub-sequence in the test sequence are complementary and anomalous deviations can be determined by combining them. Several methods are evaluated for aggregating these measures. These include a scalarized of the typicality and atypicality score, a 2-dimensional (typicality and atypicality) score, and a high-dimensional score.

Read the paper · More papers on PaperTik