Time sequence summarization: theory and applications

Quang-Khai Pham · HAL (Le Centre pour la Communication Scientifique Directe) · 2022

Domains such as medicine, the WWW or finance generate on a daily basis massive amounts of data represented as a collection of time sequence of events; Each event is described as a set of descriptors taken from various descriptive domains and associated to a time of occurrence. These archives represent valuable sources of insight for analysts to browse, analyze and discover golden nuggets of knowledge. However, discovering knowledge for such domains is a challenging task since it requires mining massive sequences where the variability of event descriptors could be very high. Recent studies show that data mining methods might need to operate on derived forms of the data including aggregate values or summaries. Knowledge extracted in such a way is called Higher Order Knowledge. In this thesis work, we propose to address this challenge and we define the concept of Time sequence summarization. The purpose is to support applications to scale on very large datasets. Time sequence summarization uses the content and temporal information of events to generate a more concise, yet informative, time sequence that can seamlessly be substituted for the original time sequence in any application. First, we propose a user-oriented technique called TSaR built in a 3-step process: Generalization, grouping and concept formation. TSaR uses taxonomies to generalize event descriptors at higher levels of abstraction. The grouping process then gathers events that are close on the timeline. Second, we reformulate the summarization task as a clustering problem to render it parameter-free. The originality resides on the specificity of the objective function to optimize. Indeed, the objective function considers both the content and the proximity of events on the timeline. We present three solutions including two greedy solutions called G-BUSS and GRASS. Finally, we explore and analyze how summaries contribute to discovering Higher Order Knowledge. We analytically characterize higher order patterns discovered from summaries and devise a methodology that uses the patterns discovered to uncover even more refined patterns. We evaluate and validate our summarization algorithms and our methodology by an extensive set of experiments on real world data extracted from Reuters's financial news archives.

Read the paper · More papers on PaperTik