Detecting Out-Of-Distribution Labels in Image Datasets With Pre-trained Networks

Susanne Wulz, Ulrich Krispel · Journal of WSCG · 2025

Ensuring the correctness of annotations in training datasets is one way to increase the trustworthiness and reliability of Machine Learning.This study aims to detect semantic shifts in datasets using Feature-Based Out-Of-Distribution and outlier detection methods, assuming Out-Of-Distribution samples are far from In-Distribution data.The experiments began with distance-based methods, such as k-Nearest Neighbours and Mahalanobis, followed by feature pyramids and dimensionality reduction techniques to address high-dimensional challenges.The results showed that the k-Nearest Neighbours detector performed robustly, achieving 100% AUROC when using ResNet50 on the Caltech-101 dataset, while the Mahalanobis detector showed unstable results with scores close to 50%.Moreover, selecting the right backbone model and feature levels, particularly low-level features from ResNet50, improved performance achieving AUROC score of 96% on the DelftBikes dataset for both k-Nearest Neighbours and Local Outlier Factor.The study highlights that k-Nearest Neighbours, Local Outlier Factor, alongside feature pyramids and dimensionality reduction constitute an effective setup for Out-of-Distribution detection, but optimal performance depends on tailored configurations across varying data conditions.

Read the paper · More papers on PaperTik