Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation

Tu Vu, Aditya Barua, Brian Lester, Daniel M. Cer, Mohit Iyyer, Noah Constant · 2022

In this paper, we explore the challenging problem of performing a generative task in a target language when labeled data is only available in English, using summarization as a case study.We assume a strict setting with no access to parallel data or machine translation and find that common transfer learning approaches struggle in this setting, as a generative multilingual model fine-tuned purely on English catastrophically forgets how to generate non-English.Given the recent rise of parameter-efficient adaptation techniques, we conduct the first investigation into how one such method, prompt tuning (Lester et al., 2021), can overcome catastrophic forgetting to enable zero-shot cross-lingual generation.Our experiments show that parameter-efficient prompt tuning provides gains over standard fine-tuning when transferring between lessrelated languages, e.g., from English to Thai.However, a significant gap still remains between these methods and fully-supervised baselines.To improve cross-lingual transfer further, we explore several approaches, including: (1) mixing in unlabeled multilingual data, and (2) explicitly factoring prompts into recombinable language and task components.Our approaches can provide further quality gains, suggesting that robust zero-shot crosslingual generation is within reach.ING and standard MODELTUNING for zero-shot cross-lingual generation (XGEN).We show that increasing model scale and decreasing tunable parameter capacity are key for overcoming catastrophic forgetting on XGEN.• We propose WIKILINGUA-0, a challenging XGEN benchmark and an associated SP-ROUGE evaluation metric, which we hope will facilitate future work evaluating multilingual summarization.• We show that mixing in unsupervised multilingual data can boost XGEN performance, and are the first to combine this approach with PROMPTTUNING.• We propose "factorized prompts", a novel approach that can also help PROMPTTUNING overcome severe catastrophic forgetting.• To facilitate future work, we release our data, pretrained models

Read the paper · More papers on PaperTik