Ranking the scores of algorithms with confidence

Adrien Foucart, Arthur Elskens, Christine Decaestecker · 2025

Evaluating algorithms (particularly in the context of a competition) typically ends with a ranking from best to worst.While this ranking is sometimes accompanied by statistical significance tests on the assessment metrics, sometimes associated with confidence intervals, the ranks are usually presented as singular values.We argue that these ranks should themselves be accompanied by confidence intervals.We investigate different methods for computing such intervals, and measure their behaviour in simulated scenarios.Our results show that we can obtain robust confidence intervals for ranks using the Iman-Davenport test and the pairwise Wilcoxon signed-rank test with Holm's correction.

Read the paper · More papers on PaperTik