Active Learning with Interpretable Predictor
Yusuke Taguchi, Keisuke Kameyama, Hideitsu Hino · 2019
Active learning is a method of constructing a useful prediction model with the minimum number of annotations or labeling for a response variable. It is widely used as a modern experimental design method, particularly for problems with high annotation cost. An appropriate reason for the selection of the next experimental setting is required for experiments with high annotation cost, for example, situations requiring large-scale experiments or long-term experiments such as agricultural examinations. In conventional active learning, it is only known that the samples that can improve the prediction accuracy of a prediction model are selected. This is not a satisfactory explanation for the selection of the next experimental setting for approving an experiment. In this paper, we propose a novel active learning algorithm with the following two models: a model to predict a response variable and a model to predict the amount of decrease in test loss. A new sample is selected using a model that predicts the amount of decrease in test loss. It is possible to provide a reason for sample selection by employing a model that can evaluate variable importance, e.g., using a random forest as a model of predicting the decrease in test loss. We applied the proposed method to multiple datasets and showed that the prediction performance of the proposed method is comparable to those of existing methods and the computational time is superior to those of existing methods. In addition, we demonstrated that it is possible to provide suitable reasons for selecting a sample in the process of active learning.