Time dependent clustering of time series
Pinto da Costa, F Joaquim, Isabel Vanessa de Assis Silva, Roberto Frías, Maria Eduarda · 2007
Pinto da Costa, Joaquim F.Faculdade de Ciencias da Universidade do Porto, Departamento de Matematica AplicadaRua do Campo Alegre,6874169-007 Porto, PortugalE-mail: [email protected], IsabelFaculdade de Engenharia da Universidade do Porto, Departamento de Engenharia CivilRua Dr. Roberto Frias4200-465 Porto, PortugalE-mail: [email protected], M. EduardaFaculdade de Ciencias da Universidade do Porto, Departamento de Matematica AplicadaRua do Campo Alegre,6874169-007 Porto, PortugalE-mail: [email protected] this work we consider the problem of clustering time series. Contrary to other works onthis topic, our main concern is to let the most important observations, for instance the most recent,have a larger weight on the analysis. This is done by defining a similarity measure between two timeseries, based on Pearson’s correlation coefficient, which uses the notion of weighted mean and weightedcovariance, where the weights increase monotonically with the time.As pointed out by Caiado et al. (2006), a fundamental problem in the clustering of time series isthe choice of a relevant metric. For us, two time series are similar to each other, and should therefore fallinto the same cluster, if their evolution over time shows similar characteristics. Consider for instancethe example in Beringer and Hullermeier (2006), where two stocks both of which continuously increasebetween 9:00 AM and 10:30 AM but then started to decrease until 11:30AM are considered similar,no matter what their absolute values are. That is to say that what interest us is not the distancebetween two time series but the distance between their “profiles”, which in our case consist in thestandardization of the two time series. We will start thus by deriving an expression for this distance.Let E = {X