Knowledge Discovery from Time Series
Tim Schlüter · Univ. Duesseldorf: Duesseldorfer Dokumenten- und Publikationsserver · 2012
Nowadays, organizations of diverse areas are collecting several kinds of data, which results in a huge bulk of data that possibly contains useful information. Since this amount of data is far too big for manual analysis, algorithms for semi-automatically discovering potential useful information within this data are developed, which is the main subject of the research area Knowledge Discovery in Databases. Among the different kinds of data, from which knowledge can be discovered, Time Series represent an especially challenging kind, since they contain interesting temporal particularities which have to be regarded separately. Analyzing time series with respect to these temporal particularities can analogously be denoted as Knowledge Discovery from Time Series, which is the main issue of this work. In order to provide the background for this thesis, we first introduce and provide a detailed review of knowledge discovery in databases in general and time series analysis in particular. After that, we introduce our contributions and integrate them into the area of Knowledge Discovery from Time Series. The first two contributions concern the subarea temporal association rule mining, which aims at analyzing transactional data with temporal information (which can be regarded as complex time series) in order to find associations within this data. Here, we introduce TARGEN, a market basket dataset generator which models several temporal coherences (which is thus ideal for testing new temporal association rule mining algorithms), and a tree-based approach for mining several kinds of temporal association rules at once. We transfer standard and temporal association rule mining techniques, which were originally designed for transactional data, to the analysis of elementary time series, and present a concrete approach for mining such standard and temporal association rules from a time series database, which for instance can be used for predicting future values of time series. In addition to that, we present two further approaches for time series analysis and prediction, which use a Hidden Markov Model basing on inter-time-serial correlations discovered by using derivative dynamic time warping and a novel motifs-based time series representation. Finally, we present two approaches applying time series analysis to concrete problems in linguistics and medicine, namely approaches for measuring text similarity and automatic sleep stages scoring.