Apprentissage profond automatisé : principes et pratique
Liu, Zhengying · theses.fr (ABES) · 2021
Automated Machine Learning (AutoML) aims at rendering the application of machine learning (ML) methods as devoid of human intervention as possible. This ambitious goal has been the object of much research and engineering since the outset of ML. The objective of this thesis is to put a formal framework around this multi-faceted problem, to benchmark existing methods, and to explore new directions. To formulate the AutoML problem in a rigorous way, we first introduce a mathematical framework that: (1) categorizes all involved algorithms into three levels (alpha, beta and gamma levels); (2) concretely defines the concept of a task (especially in a supervised learning setting); (3) formally defines HPO and meta-learning; (4) introduces an any-time learning metric that allows to evaluate learning algorithms by not only their accuracy but also their learning speed, which is crucial in settings such as hyperparameter optimization (including neural architecture search) or meta-learning. This mathematical framework unifies different sub-fields of ML (e.g. transfer learning, meta-learning, ensemble learning), allows us to systematically classify methods, and provides us with formal tools to facilitate theoretical developments (e.g. the link to the No Free Lunch theorems) and future empirical research. In particular, it serves as the theoretical basis of a series of challenges that we organized. Indeed, our principal methodological approach to tackle AutoML with Deep Learning has been to set up an extensive benchmark, in the context of a challenge series on Automated Deep Learning (AutoDL), co-organized with ChaLearn, Google, and 4Paradigm. These challenges provide a benchmark suite of baseline AutoML solutions with a repository of around 100 datasets, over half of which are released as public datasets to enable research on meta-learning. At the end of these challenges, we carried out extensive post-challenge analyses which revealed that: (1) Winning solutions generalize to new unseen datasets, which validates progress towards universal AutoML solution; (2) Despite our effort to encourage generic solutions, the participants adopted specific workflows for each modality; (3) Any-time learning was addressed successfully, without sacrificing final performance; (4) Although some solutions improved over the provided baseline, it strongly influenced many; (5) Deep learning solutions dominated, but Neural Architecture Search was impractical within the time budget imposed; (6) Ablation studies revealed the importance of meta-learning, ensembling, and efficient data loading, while data-augmentation is not critical. All code and data are available at autodl.chalearn.org. Besides the introduction of a novel general formulation of the AutoML problem, setting up and analyzing the AutoDL challenge, the contributions of this thesis include: (1) Developing our own solutions to the problems we posed to the participants. Our work GramNAS tackles the neural architecture search (NAS) problem by using a formal grammar to encode neural architectures. Two alternative search strategies have been experimentally investigated: one based on Monte-Carlo Tree Search (MCTS), which achieves 94% accuracy on CIFAR-10 dataset, and another one based on an evolutionary algorithm which beats state-of-the-art packages AutoGluon and AutoPytorch on 4 large well-known datasets; (2) Laying the basis for a future challenge on meta-learning. The AutoDL challenge series revealed the importance of meta-learning but the challenge setting did not evaluate meta-learning properly. With an intern, we experiment with various meta-learning challenge protocols; (3) Making several theoretical contributions. During the course of this thesis, several collaborations were entered to tackle problems of transfer learning and expressiveness of neural networks. Investigations on the Universal Approximation Theorem helped us understand theoretical guarantee behind Deep Learning systems we deploy.