A Double Algorithm of Web Usage Mining Based on Sequence Number
Gang Fang, Jiale Wang, Ying Hong, Jiang Xiong · 2009
Web usage mining is an application of data mining technology to mining the data of the Web server log files. It can discover these session patterns of user and some kinds of correlations between these Web pages. Web usage mining provides the support for the Web site design, providing personalization server and other business making decision. There are some session patterns saved in Web server log files, page attribute of which is Boolean quantity. In order to improve efficiency of presented algorithms and reduce the time of scanning database, and so aiming to these characters, this paper proposes a double algorithm of Web usage mining based on sequence number, which is suitable for mining any session patterns. The algorithm turns session pattern of user into binary, and then uses up and down search strategy to double generate candidate frequent itemsets. The algorithm computes support by sequence number dimension in order to scan once session pattern of user, which is different from traditional double search mining algorithm. And the efficiency of Web usage mining is efficiently improved because of this way. The experiment indicates that the efficiency is faster and more efficient than presented similar algorithms.