A Simple, Effective and Extendible Approach to Deep Multi-task Learning
Yang Gao, Yifan Li, Yu Lin, Hemeng Tao, Latifur R. Khan · 2020
Existing solutions to multi-task learning typically rely on manually enumerating multiple network architectures to find the optimal structure, which incurs a heavy design workload. In addition, extending these models to new tasks is difficult in many cases, since it often requires a network redesign for achieving the best performance. To overcome these limitations, in this paper, we propose a novel principle for multitask learning, which focuses on the learning process itself. For each task, we consider its feature as a state, and treat the learning problem as a state transformation process, which is driven by a task-specific gradient. Specifically, we introduce a Gradient Modification Unit (GMU), which consists of a Gradient Estimation Network (GEN) (shared among all tasks) and multiple tiny task-specific gradient correctors (one for each task). At each iteration, for any task T, the corresponding gradient corrector modifies the gradient estimated by the GEN to produce the desired task-specific gradient tensor that is applied to update the state of task T. Our solution has several benefits including automatic end-to-end learning of the optimal transformation process, simplification of network design and high expansibility to new tasks. We demonstrate the superiority of our approach over existing solutions on a variety of datasets, across both image classification and image retrieval tasks.