Chasing Shadows: Solving Deepfake Detection Benchmarks Using Irrelevant Features Only

Ying Xu, Philipp Terhörst, Marius Pedersen, Kiran Bylappa Raja · 2025

The emergence of Deepfake technology poses significant threats, particularly regarding misinformation and privacy. To mitigate these threats, Deepfake benchmarks play an important role in developing and testing reliable Deepfake detection algorithms. Consequently, it is crucial that these benchmarks do not possess serious biases that hinder the robustness and generalizability of Deepfake detectors during training and subsequently distort their true reliability during operation. This work investigates inherent biases in various Deepfake detection benchmark datasets by training simple classification models based on soft-biometric facial properties that do not contain Deepfake-related clues, i.e., decoy features. These mirage models reach up to 87.42% (balanced) accuracy on benchmark datasets using irrelevant decoy features alone for this task. As large parts of the performance of state-of-the-art models could also be achieved through exploiting benchmark biases, this raises the question of the unbiased performance of Deepfake detectors and their general reliability. Our analysis includes various Deepfake detection benchmarks and analyzes soft-biometric properties in determining their contribution to “solving” these benchmarks. Our findings underscore the need for more unbiased benchmarks beyond simply balancing demographic groups to enable future work on developing reliable solutions.

Read the paper · More papers on PaperTik