Enhancing Data Quality with Label Noise Detection

Wanwan Zheng · 2023

Data is a key component of machine learning performance in general. In particular, if the data has been mistakenly labeled, called label noise, the regularity of the data is difficult to determine, which may result in a reduction in inference accuracy and biases in the interpretation of the results. In order to detect such label noise, classifier-based methods are currently the most commonly used. However, there is a lack of consistency in this kind of method, since noise detection results vary depending on the classifier used. Additionally, it is computationally intensive, resulting in a delay in processing. A noise detection method was proposed in this study, whose performance is independent of classifiers, and the removal of identified noises led to the highest overall classification accuracy.

Read the paper · More papers on PaperTik