On Ensembles, I-Optimality, and Active Learning
William D. Heavlin · Journal of Statistical Theory and Practice · 2021
Abstract We consider the active learning problem for a supervised learning model: That is, after training a black box model on a given dataset, we determine which (large batch of) unlabeled candidates to label in order to improve the model further. We concentrate on the large batch case, because this is most aligned with most machine learning applications, and because it is more theoretically rich. Our approach blends three ideas: (1) We quantify model uncertainty with jackknife-like 50-per cent sub-samples (“half-samples”). (2) To select whichnofCcandidates to label, we consider (a rank- $$(M-1)$$ (M-1) estimate of) the associated $$C\times C$$ C×C prediction covariance matrix, which has good properties. (3) Our algorithm works only indirectly with this covariance matrix, using a linear-in-Cobject. We illustrate by fitting a deep neural network to about 20 percent of the CIFAR-10 image dataset. The statistical efficiency we achieve is $$3\times$$ 3× random selection.