Ranking the scores of algorithms with confidence
Adrien Foucart, Arthur Elskens, Christine Decaestecker · 2025
Evaluating algorithms (particularly in the context of a competition) typically ends with a ranking from best to worst.While this ranking is sometimes accompanied by statistical significance tests on the assessment metrics, sometimes associated with confidence intervals, the ranks are usually presented as singular values.We argue that these ranks should themselves be accompanied by confidence intervals.We investigate different methods for computing such intervals, and measure their behaviour in simulated scenarios.Our results show that we can obtain robust confidence intervals for ranks using the Iman-Davenport test and the pairwise Wilcoxon signed-rank test with Holm's correction.