Are We Making Progress? -An Analysis of NTCIR QAC1 and 2-
Masako Nomoto, Yoshio Fukushige, Mitsuhiro Sato, Hiroyuki Suzuki · NTCIR · 2004
This paper tackles issue on comparing evaluation results using multiple QA test collections(NTCIR QAC1 and 2). We identify two features that have moderate correlation with the performance of systems in QAC1 and 2 and evaluate the diculty of the two test collections using the features. Answer categories of questions also affect the performance of systems. The evaluation results suggest that QAC2 seems to be easier than QAC1 in terms of the features, and we are making progress at least for some categories. We make a proposal for the future QAC tasks, as regards to the data needed for evaluation using multiple test collections.