THUIR at TREC 2003: Novelty, Robust and Web *

Min Zhang, Chuan Lin, Yiqun Liu, Leo Zhao, Shaoping Ma · 2003

describing in following sections, respectively. A new IR system named TMiner has been built on which all experiments have been performed. In the system, Primary Feature Model (PFM) [1] has been proposed and combined with BM2500 term weighting [2] , which led to encouraging results. Word-pair searching has also been performed and improves system precision. Both approaches are described in robust experiments (section 2), and they were also used in web track experiments. 1. Novelty track Our research on this year's novelty track mainly focused on four aspects: (1) unsupervised relevance judgment; (2) efficient sentence redundancy computing; (3) supervised sentence classification; (4) supervised redundancy threshold learning. 1.1. Unsupervised relevance judgment The work of finding relevant information is useful for task1 and task3. Since words mismatch problem is dominant in sentence comparison, three kinds of approaches have been carried out in unsupervised relevance judgment to solve the problem as following. (1) Query Expansion (QE) using WordNet synonymy and hyponomy, and Dr Lin Dekang’s dictionary of term dependency [3];

Read the paper · More papers on PaperTik