NoSQL Environments and Big Data Analytics for Time Series

Ciprian‐Octavian Truică, Elena‐Simona Apostol · 2021

Times Series data are used to correlate and predict future events from sequential timestamped observations. And, with the new paradigms of information processing in the era of Big Data and with the shift in data modeling that arose with the development of NoSQL technologies, the research and industry communities alike started to look into new ways to manage efficiently and analyze accurately this type of data. In the research community, two intertwined directions have emerged: Time Series management, supported by the data management community, and Time Series analysis, backed by the machine learning and statistics communities. From the Time Series management perspective, the main objectives are: (i) optimal Time Series data processing and storing, and (ii) efficient Time Series data querying and retrieving. While, the Time Series Analysis research directions try to address and solve the following problems (i) identifying the nature of the phenomenon represented by the sequence of observations, (ii) predicting future values of the Time Series variable, (iii) outlier detection of extreme values that deviate from other observations, and (iv) change point detection for understanding trends and seasonality of Time Series. In this chapter, we present an overview of NoSQL Time Series databases and Big Data machine learning techniques, and we analyze their strengths and weaknesses to make an informed decision when choosing the best solution for processing, storing, managing, and analyzing Time Series.

Read the paper · More papers on PaperTik