Using Grammatical Relations, Answer Frequencies and the World Wide Web for TREC Question Answering.
Sabine Buchholz · 2001
This year, we participated for the rst time in TREC, and entered two runs for the main task of the TREC 2001 question answering track. Both runs use a simple baseline component implemented especially for TREC, and a high-level NLP component (called Shapaqa) that uses various NLP tools developed earlier by our group. Shapaqa imposes many linguistic constraints on potential answers strings which results in not so many answers being found but those that are found have a reasonably high precision. The dierence between the two runs is that the rst applies Shapaqa to the TREC document collection directly, whereas the second one uses it on the World Wide Web (WWW). Answers found there are then mapped back to the TREC collection. The rst run achieved a MRR of 0.122 under the strict evaluation (and 0.128 lenient), the second one 0.210 (0.234). We argue that the better performance is due to the much larger number of documents that Shapaqa-WWW's answers are based on.