Evaluation of Dataset Distribution and Label Quality for Autonomous Driving System

Sijia Li, Yong Fan, Yue Ma, Ya Pan · 2021 IEEE 21st International Conference on Software Quality, Reliability and Security Companion (QRS-C) · 2021

Nowadays, some well-known open source datasets used to train models in autonomous driving domain, which still have some data quality problems, and greatly affect the accuracy of training results. Therefore, it is meaningful to study the influence of dataset's quality. According to ISO/IEC 25012:2008(E), this paper mainly evaluates the dataset distribution and the dataset label quality. Dataset distribution is the equilibrium between the features of the dataset, and the dataset label quality includes missing, duplicate and inconsistent label. We used the noise data learning model O2U-Net [1] to see the impact on accuracy by simulating dataset quality problems, such as changing the equilibrium and making noise labels. The experimental results show that adding “car” and other common features in automatic driving filed, had a higher accuracy in the model, while missing label and inconsistent label have the most serious impact on the accuracy of the model. In terms of the determination of dataset quality, more attention should be paid to the equilibrium of interested feature in the autonomous driving field, as well as the problems of missing label and inconsistency label. Finally, this paper integrates a data quality testing platform, which can well show the quality problems of autonomous driving datasets.

Read the paper · More papers on PaperTik