Exploring the Limits of Epistemic Uncertainty Quantification in Low-Shot Settings

Matías Valdenegro-Toro · 2021

Motivation: Bayesian Deep Learning promises good uncertainty estimates, but methods often rely in approximations, and real-world datasets have issues not present in academic benchmarks (like CIFAR10, Fashion MNIST, ImageNet, etc), such as low number of samples. Evaluating the quality of output uncertainty is difficult as there are no labels. In this paper we evaluate uncertainty quantification methods as the size of the training set is varied, to simulate real-world datasets. Approach: We take random subsamples of CIFAR10 and Fashion MNIST training sets and train several uncertainty methods (7 in total), evaluating on the corresponding test set. We measure accuracy, expected calibration error, entropy, maximum probability, and out of distribution detection AUC (with SVHN and MNIST), as the training set size is varied. Contributions: We compare uncertainty methods across different training set sizes, showing that confidences do not accurately portray model uncertainty. We show that ECE and OOD detection degrades with small training sets. We provide evidence for practitioners to select uncertainty methods and give future research directions.

Read the paper · More papers on PaperTik