Subsequence time series (STS) clustering techniques for meaningful pattern discovery

Kadir A. Peker · 2006

Subsequence time-series (STS) clustering is one way of finding significant patterns in time series. A window of size w is shifted over the series to produce a w dimensional sequence and clustering is applied. Recently, it has been demonstrated that subsequence time-series (STS) clustering often produces meaningless results. We demonstrate the problem on a synthetic dataset in a number of scenarios and provide a deeper insight to the sources of the issue. More specifically, we examine the effect of trivial matches, and the proliferation of patterns under a sliding window. Although STS clustering generates meaningless cluster averages, we show that cluster cores can be used to discover real patterns. We also demonstrate the effectiveness of using a high number of clusters. First approach is using a moderately high number where the results are visually inspected. Second approach is a two stage clustering method where a very high number of clusters are generated in the first stage. Then a minimum shift distance is used to align and cluster patterns. The original patterns in our synthetic dataset are successfully recovered using this novel method.

Read the paper · More papers on PaperTik