Finding Label and Model Errors in Perception Data With Learned Observation Assertions
Daniel D. Kang, Nikos Aréchiga, Sudeep Pillai, Peter D. Bailis, Matei Zaharia · Proceedings of the 2022 International Conference on Management of Data · 2022
ML is being deployed in complex, real-world scenarios where errors have impactful consequences. In these systems, thorough testing of the ML pipelines is critical. A key component in ML deployment pipelines is the curation of labeled training data. Common practice in the ML literature assumes that labels are the ground truth. However, in our experience in a large autonomous vehicle development center, we have found that vendors can often provide erroneous labels, which can lead to downstream safety risks in trained models.