XPrompt: Exploring the Extreme of Prompt Tuning

Fang Ma, Chen Zhang, Lei Ren, Jingang Wang, Qifan Wang, Wei Wu, Xiaojun Quan, Dawei Song · 2022

Prompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner.While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performance gap between prompt tuning and fine-tuning for models of moderate and small scales (typically less than 11B parameters).In this paper, we empirically show that the trained prompt tokens can have a negative impact on a downstream task and thus degrade its performance.To bridge the gap, we propose a novel PROMPT tuning model with an eXtremely small scale (XPROMPT) under the regime of lottery tickets hypothesis.Specifically, XPROMPT eliminates the negative prompt tokens at different granularity levels through a hierarchical structured pruning, yielding a more parameter-efficient prompt yet with a competitive performance.Comprehensive experiments are carried out on the SuperGLUE tasks, and the results indicate that XPROMPT is able to close the performance gap at smaller model scales. 1 Recently, Prompt-Tuning (Lester et al., 2021; Liu et al., 2021b) has been proposed to address this issue by prepending a soft prompt to the input and only updating the parameters of prompt tokens during tuning.Prompt-Tuning provides a parameter-efficient alternative to fine-tuning, since the scale of the soft prompt is tens of thousand smaller.It is also conceptually simpler and more flexible than other parameter-efficient tuning methods (such as Adapters), that require intrusive modifications to transformer layers (Houlsby et al., 2019;Guo et al., 2021).Using fewer tunable parameters, prompt tuning achieves competitive performance to fine-tuning with the increase of the model scale.However, there is still a large performance gap between prompt tuning and fine-tuning for models of smaller scales (as shown in Figure 1).This paper aims to fill the gap, from the perspective of the lottery tickets hypothesis (LTH) (Frankle and Carbin, 2019).We are motivated by an observation that, on a specific task, not all prompt tokens contribute equally to the task performance, while certain prompt tokens may even bring a negative

Read the paper · More papers on PaperTik