Visualization of Semantically Meaningful Representations of Large Multi-dimensional Time Series on Supercomputers
Gabrielle Poerwawinata · 2020
Visualizing high-dimensional data is a cornerstone in scientific data mining. The meaningful representation of knowledge discovery in time series helps the expert to understand their data behavior. In high-dimensional data, visualization often becomes more prominent as it can help the user in identifying the patterns, classifications, and relationships within contributing variables. In time series analysis, Euclidean distance based time-series similarity plays a big role. A novel method based on Euclidean distance so-called matrix profile, has demonstrated use in pattern analysis, clustering, rule discovery. Its capability to be computed in parallel has distinguished its performance compare to other techniques. It enables us to work with many sensors and very long time-series data. Due to the irregularities of the sensors data, we use an unsupervised approach to infers the feature with no specific application domain. Therefore, the developed model should be able to be applied for any kind of time series data. Furthermore, additional study is included to infer the background of the data source. In many data investigations, the background study helps in developing a hierarchical systematic which guides the user to narrow down the exploration space for a certain aspect of analysis. After information extraction, visualization as the terminal of this study: research on insightful ways to show time series similarity, segmentation, and data clustering has to be performed. As there are many ways to visualize the data in design, information delivery, and engagement with the users, we limit the research work on the first two parts. In a small section, we also review the opportunities to create a more interactive way of visualization.