Coverage-Driven Synthetic Data Generation for Machine Learning Assurance
Manuel Hirschle, Dmitrii Kirov, Rosario Aievola, Jürgen Adamy · 2024
Data quality plays a paramount role in the assurance of safety-critical machine learning enabled systems. One of the key requirements is that training and test datasets must be complete, that is, their elements (e.g., data points, images) must sufficiently cover the space of the operational design domain for the intended application. Use of synthetic data is one of the means for improving data completeness. In this paper, we present extensions for our previously proposed scenario-based method for generating such data. Specifically, we propose and benchmark several quantitative metrics for measuring the coverage of parameter spaces. We then integrate these metrics into an iterative workflow for sampling scenarios that are fed into high-fidelity simulation tools to produce synthetic data. Such coverage-driven workflow can guide the data generation process by identifying uncovered regions of the input space and timely filling them, thus maximizing coverage with less number of samples compared to many conventional sampling methods. This is especially valuable for generating “expensive” data, such as images. We demonstrate our approach on a maritime search and rescue case study by producing image data for an object detection neural network.