Efficient Fast Updated Frequent Pattern tree algorithm and its parallel implementation
Detao Lv, Bo Fu, Xiao Sun, Hang Qiu, Xiaobing Liu, Yanlong Zhang · 2017
With the rapid expansion of data, incremental update mining for associated rule learning in large scale volume of data processing is concerned. The support count or support ratio are often used as a threshold for the associated rule learning. The support ratio is sensitive to the threshold, especially, adapted to build anomaly detection models learned with unbalanced samples. It might result in missing anomalies or data failures. However, the support count with the characteristic of good explanatory is easy to use. Therefore, the support ratio always needs to be converted into the support count. In order to mine large-scale frequent patterns, based on the FUFP algorithm, an efficient algorithm of incremental frequent item mining is proposed in this paper, named EFUFP (Efficient Fast Updated Frequent Pattern) tree algorithm. The minimum support count and a small item cache are used to effectively detect unbalanced abnormal data and improve the efficiency of query access. The parallelization of EFUFP is implemented on the Spark environment. The experimental results show that the EFUFP algorithm is suitable for the current large scale frequent item mining.