Active Learning for Regression With Correlation Matching and Labeling Error Suppression

Xiaohua Li, Jian De Zheng · IEEE Signal Processing Letters · 2016

In this letter, we develop an active learning algorithm to optimize the selection of training data for robust linear regression. This algorithm selects training data based on the principle of correlation matching between the training dataset and the overall data pool. Considering the inevitable and potentially heavy human labeling errors, we model the probability of labeling errors based on the item response theory (IRT) and develop data screening techniques to control the error sparsity. Compressive sensing theory is then exploited for human labeling error suppression. This algorithm is robust even in the case of short training dataset with nonsparse labeling errors. Its performance is verified by simulations with both artificial data and real benchmark data. Experiments are also conducted to demonstrate the validity of the IRT-based human labeling error model and the superior performance of the algorithm in practical applications.

Read the paper · More papers on PaperTik