Parting With Illusions About Synthetic Data

Daniel Pototzky, Azhar Sultan, Lars Schmidt-Thieme · 2022

Lack of labeled data is an omnipresent issue in deep learning for computer vision. Generating synthetic data seems to offer a simple solution for this problem. Once a data generator is set up, arbitrary amounts of fully-labeled data can be created. This data is then used to train a neural network which finally solves a problem on real data, e.g. by detecting cars.We argue that leveraging synthetic data like that rarely works in practice. Synthetic images often have a significant domain gap to real data, leading to reduced performance on the target domain. In experiments on several synthetic-to-real benchmarks including Sim10k to CityScapes, we show that state-of-the-art domain adaptation methods trained on thousands of synthetic images are usually outperformed by ordinary supervised learning on 14 to 70 images from the target domain.

Read the paper · More papers on PaperTik