tsmp: An R Package for Time Series with Matrix Profile
Francisco Bischoff, Pedro,Pereira Rodrigues · The R Journal · 2020
This article describes tsmp, an R package that implements the MP concept for TS.The tsmp package is a toolkit that allows all-pairs similarity joins, motif, discords and chains discovery, semantic segmentation, etc.Here we describe how the tsmp package may be used by showing some of the use-cases from the original articles and evaluate the algorithm speed in the R environment.This package can be downloaded at https://CRAN.R-project.org/package=tsmp. Introduction: time series data miningA TS is a sequence of real-valued numbers indexed in time order.Usually, this sequence is taken in a regular period of time, which will be assumed to be true in this context.The interests in TS data mining have been growing along with the increase in available computational power.This kind of data is easily obtained from sensors (e.g., ECG), (ir)regular registered data (e.g., weekly sales, stock prices, brachial blood pressure).Even other kinds of data can be converted to TS format, such as shapes (Wei et al., 2006) and DNA sequences (Shieh and Keogh, 2008).TS are generally large, high dimensional and continuously updated which requires algorithms fast enough in order to be meaningful.Besides, unlike other kinds of data, which usually have exact answers, TS are usually analysed in an approximated fashion.These characteristics have been challenging researchers to find faster and more accurate methods to retrieve meaningful information from TS.This required one or more of these methods: dimensionality reduction, constraints, domain knowledge, parameter tweaks.Only afterwards could the data mining tasks be applied in feasable time.Typical tasks include motif and discord discovery, subsequence matching, semantic segmentation, rule discovery, similarity search, anomaly detection, clustering, classification, indexing, etc. (Fu, 2011).This paper describes the tsmp package (Bischoff, 2018) which uses a novel approach to TS data mining: the MP Yeh et al. (2017b), which is based on the APSS (also known as similarity join). The APSS' task is to, given a collection of data objects, retrieve the nearest neighbour for each object. The remaining part of this paper is organised as follows: In Section 2.2 we describe the reasoning behind the MP, in Section 2.3 we present the tsmp package with examples, in Section 2.4 we compare the performance of the R implementation, and in Section 2.5 we conclude with a brief discussion.