Comparing Active Learning with Random Selection when Building Predictive Models
Fang Wu · 2021
In order to study whether active machine learning would provide more efficient construction of a predictive model than random selection of experiments, this paper describes an experiment about comparing the accuracy of the models built by these two algorithms. The experiment is based on a data set which describes the diagnosis of different cancer cells. In the experiment, the author established an active learning algorithm and a random selection algorithm, and applied them to a variety of common predictive models. The evaluation process, which includes cross-validation and learning curve, compares these two algorithms’ accuracy under different amounts of training data. It is found that in most cases, the active learning algorithm can obtain higher accuracy with less data compared with the random selection algorithm.