Issues in reuse of evaluation data and metrics
Donna K. Harman · 2003
There has been a major explosion in the number of large-scale evaluations in NLP, along with an increasing interest in these evaluations by an expanding number of research groups. An example of this is the growth of the question-answering evalation in TREC, starting with 20 participating groups in 1999 and now attracting more than 35 groups. Many of these groups are newcomers to these types of evaluations and sometimes also newcomers to research in a given area.