GUCAS at TREC 2011 Microblog Track.
Xin Zhang, Kai Hui, Ben He, Tiejian Luo · Text REtrieval Conference · 2011
The aim of GUCAS's participation in the Microblog track this year is to evaluate the eectiveness of probabilistic retrieval mod- els in combination with various sources of evidence for relevance in the context of the Twitter corpus. In our ocial runs, we use the PL2F eld-based model as the baseline, on top of which query expansion is also applied. In addition, a supplement model combining recency, au- thority and URL length is developed to retrieve authoritative and timely tweets. Finally, a language lter is used to remove non-English tweets. Our experimental results show that the language lter and URL length lter can benet the most the retrieval eectiveness. In the following-up experiments, it demonstrates that the results applying the basic mod- els improve siginicantly after removing the retweets in the preliminary results.