Relation Classification Model Performance as a Function of Dataset Label Noise

Akshay Parekh, Ashish Prabhu Anand, Amit C. Awekar · 2023

The central question that this paper address is: How does the performance of a supervised model vary with respect to the label noise in the testing and training dataset? Answering this question is crucial for many real-world applications for two main reasons. First, most datasets used for training large supervised models such as Deep Neural Networks (DNNs) are crowd-sourced. Such datasets are known to have significant label noise due to multiple factors such as the complexity of the annotation task and low wages for crowd-sourced annotators. Second, these crowd-sourced datasets are large. Reannotating such large datasets is a time-consuming and costly process. Our work aims to understand the relation between the cost of data reannotation and the performance of supervised models. Understanding this relationship can be helpful while planning for data reannotation.

Read the paper · More papers on PaperTik