An improved pagerank algorithm based on fuzzy C-means clustering and information entropy
Wenbo Zheng, Shaocong Mo, Pengfei Phil Duan, Xiaotian Jin · 2017
This paper proposes an improvement to the PageRank algorithm. Most existing PageRank algorithms expect a strong correlation among consecutively accessed webpages, which in reality should be a fuzzy relationship when a user accesses pages on an arbitrarily basis. We mine data from search-behavior logs by analyzing chronological sequential patterns, and cluster all webpages using fuzzy C clustering. The weight of each cluster is identified with information entropy, which is then used to adjust the average weight. A sample of 1 million pages is used for testing. Compared with traditional PageRank, the new algorithm decreases search time by 34.83% and increases search accuracy by 41.88%; when compared with HITS, the improvements are 31.82% and 64.04% respectively.