Representative sequence selection in unsupervised anomaly detection using spectrum kernel with theoretical parameter setting
Stefan Jan Skudlarek, H. Yamamoto · 2010
Unsupervised anomaly detection is an important topic of data mining research, especially with respect to non-numerical sequence data. However, the majority of previous algorithms features empirical parameter selection. The contribution of this study is twofold: First, we show how the Akaike Information Criterion can be used to set the parameter of the spectrum kernel. Second, a distance-based algorithm for one-class unsupervised anomaly detection is presented. The algorithm uses the distance matrix of the data to select a sequence representative of the normal class by means of robust statistics. The proposed algorithm is applied to two kinds of sequence data, showing its suitability.