Performance evaluation of top-k sequential mining methods on synthetic and real datasets
Asima Jamil, Abdus Salam, Farhat Amin · International Journal of Advanced Computer Research · 2017
Data mining techniques are used for discovering unknown hidden patterns and to predict useful information to increase the profit of various organizations.Data mining is the extraction of implicit, previously unknown, and potentially useful information from data [1].It involves methods at the intersection of artificial intelligence, machine learning, statistics, and database systems [2].The methods include classification, clustering, prediction, association rule mining, sequential pattern mining.Association rule mining discovers relationships between items in a dataset irrespective of time, whereas sequential pattern mining (SPAM) considers time or order of transactions.SPM is used in various fields such as biological sequence analysis, web log click streams, medical treatment (e.g., symptoms and diseases) [3-7].*Author for correspondence Natural disasters (e.g., earthquakes), science and engineering process, serial crime solving, telephone calling patterns, and customer purchase behaviour analysis [4].The basic idea of SPM was first introduced by [5] for the problem of customer purchase sequence, as follows: "Given a set consist of a number of sequences, where each sequence consists of a list of events or transactions and each event consists of a set of items, with a given a minimum support threshold, SPM is required to find all Sequential patterns (SPs), that is, the subsequence's whose occurrence frequency in the set of sequences is greater than minimum support threshold."It is computationally complex and challenging because such mining may create large number of candidate sequences or in other words intermediate subsequence's.Since the amount of the processed data in mining SP (sequential pattern) tends to be Research