Evaluating Face Recognition Performance on Synthetic Data: A Comprehensive Analysis of Methodologies and Benchmarks
Ángela Sánchez-Pérez, Enrique Mas-Candela, Jorge Calvo-Zaragoza · 2024
Synthetic data has become increasingly important for biometrics as an alternative to real data, thereby preventing issues associated with data protection regulations. In the realm of face recognition, synthetic generation methods are emerging to provide useful databases for training. However, a standardized protocol to compare the performance of these methods is lacking, along with an assessment framework for evaluating the performance of images generated when training face recognition models. In this paper, we report the outcomes of our efforts to assess the performance of face recognition when trained with synthetic data. We utilized 5 face recognition models and 7 synthetic datasets (plus 1 real dataset as a reference). Furthermore, for each training, four strategies for model selection—an issue typically neglected—were considered, to study their influence on final performance. Each case was evaluated under known face recognition benchmarks, with different conditions. Our results provide valuable conclusions regarding the influence of each part of the workflow, empirically confirming both known takeaways and unveiling underexplored aspects, notably the importance of model selection.