Adaptive Risk Minimization: A Meta-Learning Approach for Tackling Group Distribution Shift
Marvin Zhang, Henrik Marklund, Nikita Dhawan, Abhishek Gupta, Sergey Levine, Chelsea Finn · arXiv (Cornell University) · 2020
A fundamental assumption of most machine learning algorithms is that the training and test data are drawn from the same underlying distribution. However, this assumption is violated in almost all practical applications: machine learning systems are regularly tested under distribution shift, due to changing temporal correlations, atypical end users, or other factors. In this work, we consider the setting where the training data are structured into groups and there may be multiple test time shifts, corresponding to new groups or group distributions. Most prior methods aim to learn a single robust model or invariant feature space to tackle this group shift. In contrast, we aim to learn models that adapt at test time to shift using unlabeled test points. Our primary contribution is to introduce the framework of adaptive risk minimization (ARM), in which models are optimized for post adaptation performance on training batches sampled from different groups, which simulate group shifts that may occur at test time. We use meta-learning to solve the ARM problem, and compared to prior methods for robustness, invariance, and adaptation, ARM methods provide consistent gains of 1-4% test accuracy on image classification problems exhibiting group shift.