Multitask Learning 1997–2024: Part II Regularization and Optimization

Xiaokang Liu, Jun Rong Yu, Yutong Dai, Yishan Shen, Jianmin Chen, Jie Hu, Jin Bo Huang, Yixin Liu, Yilong Yin, Vinod Namboodiri, Brian D. Davison, Jason H. Moore, Yong Chen · Harvard Data Science Review · 2025

In Part I of this survey, we introduced the fundamentals of multitask learning (MTL), including its definition, significance, and core operational mechanisms. This Part II builds on that foundation by summarizing existing knowledge and reviewing key methodologies that underpin MTL, with a focus on regularization and optimization techniques that facilitate effective information sharing across tasks. We begin by exploring a variety of regularization approaches, including sparsity-inducing feature selection methods, low-rank methods, weight matrix decomposition techniques, prior-sharing strategies, and task clustering methods, each offering a unique lens through which to capture task similarities and disentangle shared from task-specific structures. Next, we delve into optimization strategies, including scalarization approaches, multiobjective optimization methods, adversarial training, and neural architecture search, which are used to align gradient directions across tasks and improve convergence. For each category of methods, we present key formulations and technical details of representative works, concluding with a summary of insights and practical takeaways. By presenting existing methods and summarizing their strengths and limitations, this part of the survey aims to provide conceptual insights that may inspire the broader application of MTL and inform future research directions.“Multi-Task Learning 1997–2024” is a three-part article. Part I, “Fundamentals” can be read here . Part III, “Applications” can be read here .

Read the paper · More papers on PaperTik