Approximate mining of consensus sequential patterns

Hye‐Chung Kum, Wei Wang, Dean F. Duncan · 2004

Sequential pattern mining finds interesting patterns in ordered lists of sets. The purpose is to find recurring patterns in a collection of sequences in which the elements are sets of items. Mining sequential patterns in large databases of such data has become an important data mining task with broad applications. For example, sequences of monthly welfare services provided over time can be mined for policy analysis. Sequential pattern mining is commonly defined as finding the complete set of frequent subsequences. However, such a conventional paradigm may suffer from inherent difficulties in mining databases with long sequences and noise. It may generate a huge number of short trivial patterns but fail to find the underlying interesting patterns. Motivated by these observations, I examine an entirely different paradigm for analyzing sequential data. In this dissertation: (1) I propose a novel paradigm, multiple alignment sequential pattern mining, to mine databases with long sequences. A paradigm, which defines patterns operationally, should ultimately lead to useful and understandable patterns. By lining up similar sequences and detecting the general trend, the multiple alignment paradigm effectively finds consensus patterns that are approximately shared by many sequences. (2) I develop an efficient and effective method. ApproxMAP (for APPROXimate Mbultiple Ablignment Pbattern mining) clusters the database into similar sequences, and then mines the consensus patterns in each cluster directly through multiple alignment. (3) I design a general evaluation method that can assess the quality of results produced by sequential pattern mining methods empirically. (4) I conduct an extensive performance study. The trend is clear and consistent. Together the consensus patterns form a succinct but comprehensive and accurate summary of the sequential data. Furthermore, ApproxMAP is robust to its input parameters, robust to noise and outliers in the data, scalable with respect to the size of the database, and outperforms the conventional paradigm. I conclude by reporting on a successful case study. I illustrate some interesting patterns, including a previously unknown practice pattern, found in the NC welfare services database. The results show that the alignment paradigm can find interesting and understandable patterns that are highly promising in many applications.

Read the paper · More papers on PaperTik