Evaluating the Evaluations: A Perspective on Benchmarks
Omar Alonso, Kenneth Church · ACM SIGIR Forum · 2024
More and more benchmarks, datasets, and evaluation tasks are becoming available. This is extremely useful for the community because it enables researchers and practitioners to test and evaluate new techniques. However, the construction, evaluation, and maintenance of data sets and benchmarks is opaque which creates problems with respect to stability and true representations. Our position is that we need to revisit how we design and implement benchmarks. The SPEC benchmark offers interesting perspectives that our community should consider. We use a data set of influential papers and resources to discuss important benchmark aspects such as realistic workloads, reliability, validity, leakage, and labeling. We conclude by proposing a list of principles for constructing evaluation benchmarks.