All patched up: effective integration of real and synthetic features into a single image for object detection

Ashley S. Dale, Lauren Christopher, William Reindl, Edwin Sanchez, Sam Brunes, Will Bickel, Jasmine Martin, Albert M. William · 2023

Synthetic data and data augmentation are commonly used to supplement a small set of real images and create a dataset with diverse features, improving the robustness of a computer vision model. However, there is no consensus on how to best leverage synthetic data for object detection tasks in a real environment. Our work answers the following questions: First, is it more important for synthetic data to have realistic backgrounds or realistic instances? Second, is pixel saliency a valid metric for explaining the relative importance of real and synthetic instances? We explore these questions by placing synthetic object instances in real environments and vice versa to create synthetically blended data, rather than a dataset with distinct real and synthetic images. Our results show that when there is good overlap between the feature domains of the real and synthetic images, there is no benefit to augmenting the training data with patches. However, given a domain gap between real and synthetic data for a given class, model performance improved ≈ 60%, with greater gains observed when increasing the number of images with real backgrounds. These results show that the relationship of synthetic to real features should be considered when constructing a dataset for domain transfer tasks.

Read the paper · More papers on PaperTik