Time-series similarity problems and well-separated geometric sets

Béla Bollobás, Gautam Das, Dimitrios Gunopulos, Heikki Mannila · 1997

Given a pair of nonidentical complex objects, defining (and determining) how similar they are to each other is a nontrivial problem. In data mining applications, one frequently needs to determine the similarity between two time series. We analyze a model of timeseries similarity that allows outliers, different scaling functions, and variable sampling rates. We present several deterministic and randomized algorithms for computing this notion of similarity. The algorithms are based on nontrivial tools and methods from computational geometry. In particular, we use properties of families of well-separated geometric sets. The randomized algorithm has provably good performance and also works extremely efficiently in practice. 1 Introduction Being able to measure the similarity between objects is a crucial issue in many data retrieval and data mining applications; see [10] for a general discussion on similarity queries. Typically, the task is to define a function Sim(X;Y ), where X and Y are...

Read the paper · More papers on PaperTik