Representative Information Retrieval Algorithm Based on PageRank Algorithm and MapReduce Model
Ling Wei, Yang Li, Yongjiang Wei · International Journal of Database Theory and Application · 2016
Representative information retrieval is widely used in public opinion management, mobile commerce and knowledge management.In the era of big data, the value of representative information extraction is particularly prominent.In order to further explore the application value of representative information extraction, this paper proposes a representative information retrieval method based on PageRank algorithm and MapReduce model (PM-Rep), research on effect of the representative coefficient λ to the extraction scale and the coverage and redundancy of the results, comparative analysis of the advantages and disadvantages of the similar methods.In the parameter experiment, as the λ increases, the scale of the extraction results increases, the coverage degree and the redundancy degree also increases.In the effect experiment, PM-Rep's coverage is significantly better than Top-k, Heuristic and Random, and the redundancy of PM-Rep is the least.In the efficiency experiment, PM-Rep takes the least time in the four methods, embodies the advantages of PM-Rep method for massive data.