Mini-Model Adaptation: Efficiently Extending Pretrained Models to New Languages via Aligned Shallow Training

Kelly Marchisio, Patrick Lewis, Yihong Chen, Mikel Artetxe · 2023

Prior work shows that it is possible to expand pretrained Masked Language Models (MLMs) to new languages by learning a new set of embeddings, while keeping the transformer Recent work on multilingual NLP has focused on pretraining (masked) language models on unlabeled corpora in multiple languages (Pires et al., 2019;Conneau et al., 2020;Xue et al., 2021).The resulting models can then be finetuned using labeled downstream data in a single language (typically English), and zero-shot transferred to the rest of the languages.While effective, existing models rarely cover more than a few dozen languages, and pretraining new models from scratch to support additional languages can be prohibitively expensive.

Read the paper · More papers on PaperTik