Creation and Testing of Synthetic Datasets for Training Road Scenes Algorithms
Khulan Khalzaa, Stephen Karungaru, Kenji Terada · 2023
Deep learning models require large amounts of data to be trained to fulfill their potential. To solve this problem, we propose a novel method for creating high-quality photorealistic synthetic training data and compare its performance to real data for object detection. We introduce the RealStreet and SynthStreet datasets, which were designed to enhance a safety analysis for road object detection. The objective of the project is to provide a useful synthetic environment for learning and building a road traffic experience by imitating the real environment as nearly as possible. This will improve the safety of road users and enable testing and planning before actual, dangerous events arise. The RealStreet data-set was collected in real-world scenarios of an urban city, while the SynthStreet closely recreates the RealStreet scenes: field of view, road objects, such as pedestrians, cyclists, vehicles, and background information buildings are matched. We study the performance and behavior of a network model trained on real and synthetic data-sets and both with various ratio mixes. Our approach is to evaluate data-set performances using a state-of-the-art method for object detection tasks in learning from synthetic data. We also compare the performance of each dataset, which is evaluated on real-world data, to determine the possibility of synthetic data-set benefits and analyze the effect of limited real-world data. However, it is important to note that synthetic data-sets may not always accurately represent the variability and complexity of real-world environments and may not generalize well to real-world scenarios. Therefore, it is important to carefully evaluate the performance of deep learning models trained on synthetic data and validate their results using real-world data.