Anomaly Detection in Sequence Data Based on MSM-H with EXPoSE
Renkang Geng, Bin Sun, Liyao Ma, Qinjun Zhao, Tao Shen · 2021
One efficient way of anomaly detection is to accurately predict future observations by using existing data sets then evaluate the anomaly degree or score of the data by comparing the residuals between the observed values and the predicted values. Therefore, reasonable prediction of future observation values and appropriate anomaly scores are the keys to the effectiveness of anomaly detection. In this work, we proposed a new method Multi-Sub-Model with Heteroscedasticity (MSM-H) modelling for pre-processing sequence data to get residuals together with EXPected Similarity Estimation (EXPoSE) for detection using the residuals. We compare the proposed pre-processed residuals with two more datasets and compare the detection methods with five more latest and proven effective detection algorithms as the anomaly detector. Experiments are carried out using real-word traffic flow data. It firstly generates multiple small models from the traffic flow sequence data, then adapted the model to the specific data of each day, and finally processed and solves the heteroscedasticity in the data to obtain relatively accurate and reasonable residuals. Anomaly detection on the residuals shows promising results and gives 41% improvement considering Area Under Curve (AUC).