University of Waterloo at TREC 2015 Microblog Track

Luchen Tan, Adam Roegiest, Charles L. A. Clarke · Text REtrieval Conference · 2015

Given a topic with title, narrative and description, we start by building a language model for the topic. The top 1000 tweets were retrieved from Twitter commercial search engine by applying the title of the topic as a query. We exploit pseudo relevance feedback technologies to estimate probability distributions of each term in the topic, then comparing these probabilities with a background distribution model. We select the highest dierent terms as our expanded query terms. We then generate a vector for each topic, the features of the vector are non-stop word title terms, selected narrative terms and query expansion terms. Dierent weights are assigned to the dierent types of terms. Since we are allowed to deliver at most 10 tweets every day, and the latency time can not exceed 100 minutes, we solve the tweet notication scenario as a multiple-choice secretary problem. Two dierent solutions were tested.

Read the paper · More papers on PaperTik