Time series classification with Bag-Of-Words approach

Domen Kavran · 2010

The amount of data generated every day increases each year and the pace is accelerating with development of Internet of Things (IoT).Gathered data solely doesn't contain much information, but with machine learning additional informations and hidden patterns can be obtained to contribute to time series analysis.General purpose and problem specific time series feature extraction methods have been developed over the years.New feature extraction approach, derived from Bag-Of-Words, is presented in this paper.Main part of the approach is obtaining a dictionary of time series segments -the so-called words.K-Means clustering is used to form a dictionary containing K words, which is then used to define a feature vector of an individual time series as a histogram of word occurances inside it.Described approach can be used for feature extraction of time series without prior knowledge of data's nature.Moreover, the approach is robust and produces good classification results.Highest accuracy of 99.96% was achieved using datasets, presented in Results.

Read the paper · More papers on PaperTik