Finding Non-Trivial Patterns and Structural Similarity in Time Series Databases

Jessica Lin · 2009

Perhaps the most commonly encountered data type is time series. Apart from the obvious problem of handling the typically massive size of time series databases—gigabytes or even terabytes are not uncommon—most classic data mining algorithms do not perform or scale well on time series data. This is mainly due to the inherent structure of the data: high dimensionality and feature correlation. These intrinsic structural characteristics, combined with the measurement-induced noises that beset real-world time series data, pose challenges that render classic data mining algorithms ineffective and inefficient. The emphasis of this talk is on the discovery of important patterns in time series data. The previous body of work in this area has been mostly concentrated on the identification of previously known patterns. The major distinction of this work is that it offers the ability to discover important, unknown patterns in an effective and automated manner. We introduced SAX (Symbolic Aggregate approXimation), the first symbolic representation of time series that allows dimensionality reduction and lower-bounding distance measures. I will discuss SAX and its utilities, including recent work on finding structural similarity on time series data

Read the paper · More papers on PaperTik