Unified fidelity metric for synthetic imagery
Katherine Plas, Justin Zhan, Jing Lin · 2026
Object detection methods often incorporate synthetic samples into their training datasets to address limitations of real data. However, different mod- els’ sensitivity to the quality and ratio of synthetic data relative to real data varies significantly. The application of a robust metric to evaluate the qual- ity of datasets before training can result in substantial savings in time and resources. Our research investigates the relationship between data quality and model performance while exploring advantageous structures of object detection mod- els. Previous studies have focused on traditional one-stage and object-prior- based detectors, while diffusion detectors have not been extensively stud- ied for their resilience to synthetic data. We evaluate the performance of anchor-based, anchor-free, one-stage, and two-stage models on synthetic datasets, analyzing how combining a vision transformer-simple feature pyra- mid (ViT/SFP) backbone and neck can enhance performance, particularly in scenarios involving significant domain gaps. This evaluation leads to a multiple linear regression function which predicts average precision based on varying proportions and qualities of synthetic data, serving as a dynamic indicator for data quality. We find that our novel application of combined quality metrics calculated across different embedded spaces yield a stronger correlation with model performance than metrics assessed in isolation. Our research also contributes to AI by demonstrating that diffusion-based object detection models show a notable resilience to domain gaps, offering promising advancements for applications in aerial imagery datasets.