Sampling learning based association rules mining algorithm

Xiaoying Xie, Ying Zhang, Yingtao Xu · 2012

The view that sampling technology could improve the efficiency of data mining significantly has been widely accepted by the research community. The key to sample in data mining is how to design a sampling strategy to get a favorable sample to execute the mining algorithm at minor cost of accuracy. In this article we propose a progressive sampling algorithm based on confusion matrix to determine the optimal sample size. The novelty of this algorithm is that it can find the appropriate sample very quickly and very accurately without executing the data mining.

Read the paper · More papers on PaperTik