Automated Selection of High-Quality Synthetic Images for Data-Driven Machine Learning: A Study on Traffic Signs
Daniela Horn, Lars A.L. Janssen, Sebastian Houben · 2021
The utilization of automatically generated image training data is a feasible way to enhance existing datasets, e.g., by strengthening underrepresented classes or by adding new lighting or weather conditions for more variety. Synthetic images can also be used to introduce entirely new classes to a given dataset. In order to maximize the positive effects of generated image data on classifier training and reduce the possible downsides of potentially problematic image samples, an automatic quality assessment of each generated image seems sensible for overall quality enhancement of the training set and, thus, of the resulting classifier. In this paper we extend our previous work on synthetic traffic sign images by assessing the quality of a fully generated dataset consisting of 215,000 traffic sign images using four different measures. According to each sample's quality, we successively reduce the size of our training set and evaluate the performance with SVM and CNN classifiers to verify the approach. The comparability of real-world and synthetic training data is investigated by contrasting several classifiers trained on generated data to our baseline w.r.t. actual misclassifications during testing.