Pattern Discovery in Time Series
Gerhard Klassen · Univ. Duesseldorf: Duesseldorfer Dokumenten- und Publikationsserver · 2021
Abstract The identification of groups in data sets, also called cluster analysis or clustering, is an important part of many analyses. Several algorithms from different research areas have already been developed for this purpose. These methods differ not only in their algorithmic procedure but also in the use of different comparison functions. In addition, many methods require the selection of one or more parameters, so that the results depend not only on the chosen method but also on the parameters selected. The question of the validity of the clusters found can only be answered, if at all, by experts in the relevant data domain. This problem affects all forms of data and can severely limit the usefulness of such an analysis. Some types of data contain additional dependencies that can be usedadvantageously in such a cluster analysis. Time series, i.e. ordered sequences of observations, represent such a class of data. In many respects they determine our everyday life, whether on stock markets, in medicine or during the Corona pandemic in form of the course of infections. If the temporal component is properly taken into account, a cluster analysis can provide previously unknown information. However, the validity of the found clusters must first be ensured in order to prevent misinterpretations. The explained problem is the motivation for CLOSE, a new method presented here, which is able to evaluate a clustering of time series. The developed evaluation is based on a novel stability measure for time series and clusters and provides a score, which makes different clusterings comparable. The circumstance that it is not only crisp clustering which is affected by the described problem, but also fuzzy clustering, led us to another method, called FCSETS, which is specialised on fuzzy clusterings. We evaluate these methods using several data sets and clustering algorithms. We also present three applications and several variants which target the detection of outliers in time series. These applications are based on the findings in FCSETS and CLOSE. Additionally we present a clustering algorithm, which is based on a derived concept of CLOSE. In an excursion chapter, we show the results of other machine learning techniques so that a comparison can be made with our applications. Our results are promising and enable users to choose a suitable clustering algorithm and the corresponding parameters without prior knowledge. ---------------------------------------------------------------------