Further Analysis of Whether Batch and User Evaluations Give the Same Results with a Question-Answering Task.
William Hersh, Andrew H. Turpin, Lynetta Sacherek, Daniel Olson, Susan Price, Benjamin K S Chan, Dale F. Kraemer · 2000
In the TREC-8 Interactive Track, our results indicated that the better performance obtained in batch searching evaluation do not translate into better performance by users in an instance recall task. This year we pursued this investigation further by performing the same experiments using the new questionanswering task adopted in the TREC-9 Interactive Track. Our results once again show that better performance in batch searching evaluation does not translate into gains for real users. A continuing unanswered question in information retrieval (IR) research is whether batch and user searching evaluations give the same results. We explored this question in the TREC-8 Interactive Track, where we found that the better results obtained in batch studies using the Okapi weighting scheme over the standard TFIDF approach did not accrue to real users for an instance recall task.[1] This work was limited by the small number of queries as well as the use of a single retrieval task, the recall of specific instances for a topic. Since the TREC-9 Interactive Track would be using a different task- questionanswering- we decided to use the same research question again with this changed task. Although we would still have a small number of queries, it would provide another IR task to assess this research question. As with the TREC-8 Interactive Track we performed three experiments. The first experiment was to identify an IR approach that achieved the best possible performance in the batch environment. In the