Optimal Data Partition for Semi-Automated Labeling

Daniel Lopresti, George Nagy · 2013

In a pattern recognition sequence consisting of alternating steps of interactive labeling, classifier training, and automated labeling (e.g., CAVIAR systems), the choice of sample size at each step affects the overall amount of human interaction necessary to label all the samples correctly. The appropriate splits depend on the error rate of the classifier as a function of the size of the training set and, perhaps surprisingly, are independent of the relative costs of interactive correction and confirmation. We model such a system and report the sequence of optimal data partitions for a representative range of parameters. 1.

Read the paper · More papers on PaperTik