Meta-Learning via Hypernetworks

Dominic Zhao, Seijin Kobayashi, João Sacramento, Johannes von Oswald · Repository for Publications and Research Data (ETH Zurich) · 2020

Recent developments in few-shot learning have shown that during fast adaption, gradient-based meta-learners mostly rely on embedding features of powerful pretrained networks.This leads us to research ways to effectively adapt features and utilize the meta-learner's full potential.Here, we demonstrate the effectiveness of hypernetworks in this context.We propose a soft weight-sharing hypernetwork architecture and show that training the hypernetwork with a variant of MAML is tightly linked to meta-learning a curvature matrix used to condition gradients during fast adaptation.We achieve similar results as state-of-art model-agnostic methods in the overparametrized case, while outperforming many MAML variants without using different optimization schemes in the compressive regime.Furthermore, we empirically show that hypernetworks do leverage the inner loop optimization for better adaptation, and analyse how they naturally try to learn the shared curvature of constructed tasks on a toy problem when using our proposed training algorithm.Recent work has shown that MAML is mostly learning general features rather than finding fastadaptable weights deep inside its model.In fact, it was demonstrated that few-shot learning in the hidden layers of a network has little to no effect on the performance of MAML [2, 3].Furthermore, powerful models trained on rich enough data and without explicit meta-learning were shown to outperform most gradient-based meta-learning methods [4].This is consistent with huge models being few-shot learners without explicit training at all [5].These previous findings point to the untapped potential of fast adaptation within the neural network as a promising area of improvement for such meta-learning methods.One promising scalable approach is to separate the model into shared meta-parameters and context parameters [6].Here, the context parameters are the only parameters updated in the inner loop.This way, the context parameters can learn task specific information and quickly adapt while the meta-parameters are used as general reusable features.Other approaches attempt to explicitly modulate the inner-loop training procedure by meta-learning learning rates or factorised preconditioning matrices [7][8][9].Here, the actual model stays untouched and additional parameters are learned only to modulate the gradient with respect to the model parameters while learning new tasks.4th Workshop on Meta-Learning at NeurIPS 2020, Vancouver, Canada.Here, we provide new insights on how hypernetworks [10,11] implicitly combine these two seemingly different approaches.When trained with a variant of MAML, we show that hypernetworks implicitly modulate the inner loop optimization and adapt hidden layer features in a task-dependent manner.More generally, we demonstrate that hypernetworks learn features that directly support fast adaptation without any hand-designed add-ons or optimization variants.We propose a specific soft weightsharing hypernetwork architecture and show that it achieves state-of-the-art results compared to other gradient-based methods.Our method performs comparable with MAML even if the number of parameters is drastically compressed.Our main contributions are as follows:• We demonstrate the effectiveness of hypernetworks for fast adaptation -both in the compressed and overparametrized regime.• By proposing a simple soft weight-sharing hypernetwork architecture, we outperform most MAML variants without any explicit add-ons or optimization algorithm changes.Furthermore, we show empirically that hypernetworks can indeed learn useful inner-loop adaptation information and are not simply learning better network features.• We show theoretically that in a simplified toy problem hypernetworks can learn to model the shared structure that underlies a family of tasks.Specifically, its parameters model a preconditioning matrix equal to the inverse of the tasks' shared curvature matrix.

Read the paper · More papers on PaperTik