Significance tests for the evaluation of ranking methods
Stefan Evert · 2004
This paper presents a statistical model that interprets the evaluation of ranking methods as a random experiment.This model predicts the variability of evaluation results, so that appropriate significance tests for the results can be derived.The paper concludes with an empirical validation of the model on a collocation extraction task.