Synthetic Dataset Generation Methods for Computer Vision Application
Matej Arlović, Davor Damjanović, Franko Hržić, Josip Balen · 2024
Deep learning models rely on datasets for training, validation, and testing. They require diverse and representative datasets to represent real-world scenarios effectively. The performance of deep learning models is heavily influenced by data quality and quantity. High-quality and diverse datasets are essential for developing robust and accurate models. However, acquiring data for datasets can pose challenges since the time and complexity associated with manual data labeling. In order to address this problem, synthetic data generated by 3D software like Unreal Engine 5, Blender, and NVIDIA Omniverse, as well as AI-generated images in Stable Diffusion and DALL-E 3, are employed. Synthetic data facilitate the rapid and straightforward generation of aggregated datasets that accurately represent real-world scenarios. The impact of synthetic data on the performance and quality of deep learning model outcomes requires further investigation.