Web mining techniques for query log analysis and expertise retrieval
Irwin King, Micahel R. Lyu, Hongbo Deng · 2009
With the large increase in the amount of information available online, rich Web data can be obtained on the Internet, such as over one trillion Web pages, millions of scientific literature, and different interactions with society, like question answers, query logs. Currently, Web mining techniques has emerged as an important research area to help Web users find their information need. In general, Web users express their information need as queries, and expect to obtain the needed information from the Web data through Web mining techniques. To better understand what users want in terms of the given query, it is very essential to analyze the query logs. On the other hand, the returned information may be Web pages, images, and other types of data. Beyond the traditional information, it would be quite interesting and important to identify relevant experts with expertise for further consulting about the query topic, which is also called expertise retrieval. The objective of this thesis is to establish automatic content analysis methods and scalable graph-based models for query log analysis and expertise retrieval. One important aspect of this thesis is therefore to develop a framework to combine the content information and the graph information with the following two purposes: 1) analyzing Web contents with graph structures, more specifically, mining query logs; and 2) identifying high-level information needs, such as expertise retrieval, behind the contents. For the first purpose, a novel entropy-biased framework is proposed for modeling bipartite graphs, which is applied to the click graph for better query representation by treating heterogeneous query-URL pairs differently and diminishing the effect of noisy links. Based on the graph information, there is a lack of constraints to make sure the final relevance of the score propagation on the graph. To tackle this problem, a general Co-HITS algorithm is developed to incorporate the bipartite graph with the content information from both sides as well as the constraints of relevance. Extensive evaluations on query log analysis demonstrate the effectiveness of the proposed models. For the second purpose, a weighted language model is proposed to aggregate the expertise of a candidate from the associated documents. The model not only considers the relevance of documents against a given query, but also incorporates important factors of the documents in the form of document priors. Moreover, an important approach is presented to boost the expertise retrieve by incorporating the content with other implicit link information through the graph-based re-ranking model. Furthermore, two community-aware strategies are developed and investigated to enhance the expertise retrieval, which are motivated by the observation that communities could provide valuable insight and distinctive information. Experimental results on the expert finding task demonstrate these methods can improve and enhance traditional the traditional expertise retrieval models with better performance.