SAX-Based Representation with Longest Common Subsequence Dissimilarity Measure for Time Series Data Classification

Mariem Taktak, Slim Triki, Anas Kamoun · 2017

In Time Series Classification (TSC) problems, parametric extension of the longest common subsequence (LCSS) using discrete derivative has proved its superiority to the classic LCSS method. Despite its simplicity, the major shortcoming of the discrete derivative approximation is its sensitivity to noise. The parametric extension of the LCSS can prevent this drawback by comparing two time series based on their overall shape instead of a point-to-point comparison. However, global shape comparison requires multiple calculations of the LCSS dissimilarity matrix. Typically, global constraints are often employed to limit the search space in the matrix during the recursive calculation of the longest matching subsequence. Recently, an extensive study on the influence of global constraints on (dis)similarity measures shows that the performance of the 1-Nearest Neighbor (1NN) classification with LCSS decrease significantly for a constraint less than 15% to 10% of the time series length. In this work we present an approximation of derivative that requires only one recursive calculation of the LCSS (dis)similarity measure to achieve 1NN classification with optimal memory requirement. Hence, we propose a combination of advanced Symbolic Aggregate approXimation (SAX) representation with LCSS between compressed sequences of symbols. Advanced SAX aims to add symbolic trend information by applying Piecewise Linear Regression. Once time series are abstracted into symbolic representation, data compressed can be used to simultaneously reduce memory and recursive computation of the length of the longest common subsequence between run-length-encoding symbols. Experiments on UCR archives dataset have shown the importance of using advanced SAX with LCSS for TSC.

Read the paper · More papers on PaperTik