Contextual performance metrics: using synthetic data and automated characterization to contextualize deep neural network performance

Joshua Haley, Brandon Kessler, Jonathan Nesper, Matthew Waldrep, Jared K. Cooper, Michael D. DeVore · 2025

Machine Learning (ML) and Artificial intelligence (AI) have increased automation potential within defense applications such as border protection, compound security, and surveillance applications. One common model type used in these systems is object detection, an approach to detecting and localizing objects of interest within images. However, challenges exist in transitioning these state-of-the-art detection models, which are intrinsically based upon the data used to tune them, into a more general test suite for verification and validation. While metrics such as missed detections, false detection rates, and receiver operator characteristic (ROC) curves are commonly used to characterize performance, they are broad metrics that only characterize average performance, losing the contextual environment impacts on model performance. A situation arises where a model may exhibit excellent precision and recall but fails its validation activities when integrated into a system due to the testing environment being skewed toward performance outliers in the model’s testing domain. In conjunction with partners, Elbit America has built upon its synthetic training data generation capability to generate training and evaluation data across various situations to address these challenges. These data are characterized using automated image characteristic methods corresponding to environment, object, and sensor conditions. The variety of conditions required is prohibitively expensive without a data generation capability. With this data characterization, a performance model is developed that relates object detection performance to the characteristics of the input. This performance model has two primary uses. First, it can identify situations where the model is currently performing poorly. In conjunction with the synthetic data generation system, this performance model can yield an automated incremental improvement system whereby a system using a limited form of “introspection” can self-generate data and train for improvement. Second, integration into probabilistic simulations and course-of-action analysis tools yields far more nuanced results and prevents unexpected model validation results, and increases trust in AI systems. Integration into systems-level simulations provides insight into the detection model’s contribution to the system’s capabilities in the field.

Read the paper · More papers on PaperTik