Generative AI synthetic data for computer vision tasks

Shannon Dutchie, Edward Ryan, Michael F. Finch, Kimberly E. Manser, James Uplinger, John Vines, Robert Nguyen · 2025

Synthetic data has historically been used to supplement and augment real data, particularly in computer vision applications. However, as new approaches to generating synthetic data emerge, it’s important to evaluate these methods against objectives (e.g. computational resource alignment, schedule requirements, or performance thresholds). Generative AI (GAI) image creation aims to mimic specified attributes without the labor- and compute-intensive efforts associated with physical simulations. This paper investigates the effectiveness of GAI models by comparing the image quality of real data with GAI-created data for the classification task. The synthetic data is created by processing real images with a Canny edge detector, and then using a LoRA (Low-Rank Adaptation) model to in-paint based on curated prompts. In our experimental design, we first establish a baseline evaluation of real images using latent space distance metrics and projections. With this baseline as a reference point, we repeat the evaluation with the LoRA-generated imagery and compare results. As an upper bound comparison, we finally evaluate the same metrics for a real but out-of-domain dataset.

Read the paper · More papers on PaperTik