Improving Relevance Prediction by Addressing Biases and Sparsity in Web Search Click Data

Qi Guo, Dmitry Lagun, Denis Savenkov, Qiaoling Liu · 2012

In this paper, we present our approach and findings in participating the 2012 Yandex Relevance Prediction Challenge. Our approach has two goals: on one hand, we aim to address four types of biases, namely, position-bias, perception-bias, query-bias, and session-bias to better interpret the clickthrough information; on the other hand, we aim to address the clickthrough sparsity by exploiting various back-off strategies. We use gradient boosted regression trees to combine the different features and model the interactions among them. Our final submission ranks 3rd (AUC 0.6635) among the prize eligible participants on the first subset of test queries, but drops to 8th (AUC 0.6536) on the second (hidden) subset, which is potentially due to over-fitting. In this paper, we also discuss our post-competition efforts in addressing this issue through crossvalidation and more careful model selection.

Read the paper · More papers on PaperTik