Creating Robust Data Sets for AI by Leveraging Ontological Structures
Lynn Vonderhaar, Timothy Elvira, Tyler Procko, Omar Ochoa · 2025
Machine Learning (ML) models are essential across numerous fields, including those where safety is paramount, e.g., autonomous driving, which is the primary focus of this paper. In many areas, the opaque nature of ML systems may be a minor issue, but in safety-critical contexts, it creates challenges in establishing trust. To fully harness ML models in these high-stakes fields, a method is needed to enhance confidence in their reliability and accuracy without the need for human oversight on every decision. This research introduces a technique aimed at boosting trust in ML models by improving the robustness and thoroughness of the training datasets. Since ML models are shaped by the data they are trained on, ensuring the dataset's completeness can potentially enhance confidence in the model's performance. To achieve this, the paper suggests leveraging both a domain-specific ontology and an image quality ontology to assess the coverage and robustness of the dataset. A case study is presented within the context of emergency road vehicles.