Detecting cycles in Complex Time Series Databases
Francesco Giordano, Maria Lucia Parrella, Marialuisa Restaino · 2008
A large part of the data stored in financial, medical and scientific databases consists in time series realizations of unknown processes. The automatic statistical mod- elling of such data may be a very hard problem when the time series show features, such as nonlinearity, local nonstationarity, high frequency, long memory and periodic components. In such a context, the aim of this paper is to analyze the problem of detecting automatically the different periodic components in the data, with particular attention to the short term components (weakly, daily and intra-daily cycles). We focuses on the analysis of real time series from a large database provided by an Italian electric company. This database may be considered in the previous sense. Time series data mining has attracted great attention in the statistical community in recent years. The automatic statistical modelling of large databases of time series may be a very hard problem when the time series show features, such as nonlinearity, nonstationarity, high frequency, long memory and periodic components. Classification and clustering of such complex objects may be particularly beneficial for the areas of model selection, intrusion detection and pattern recognition. As a consequence, time series clustering represents a first mandatory step in temporal data mining research. In this paper we consider the case when the time series show periodic components of different frequency. Most of the data generated in the financial, medical, biological and other fields present such features. A clustering procedure useful in such a context has been proposed in Giordano, La Rocca and Parrella (2008). It is based on the detection of the dominant frequencies explaining large portions of variation in the data through a spectral analysis of the time series. Like every clustering methods, also the procedure proposed in Giordano, La Rocca and Parrella (2008) depends on some sort of tuning parameters, which directly influence the results of the classification. The purpose of this paper is to perform a sensitivity analysis on such parameters, and to control if there is a particular behaviour in the results which can help in the determination of the optimal partition and in the dynamical detection of cycles in the database. In the third section we propose a method for the selection of the best partition for the clustering procedure.