Embedding Structure Matters: Comparing Methods to Adapt Multilingual Vocabularies to New Languages
C.m. Downey, Terra Blevins, Nora Goldfine, Shane Steinert‐Threlkeld · 2023
Pre-trained multilingual language models underpin a large portion of modern NLP tools outside of English.A strong baseline for specializing these models for specific languages is Language-Adaptive Pre-Training (LAPT).However, retaining a large cross-lingual vocabulary and embedding matrix comes at considerable excess computational cost during adaptation.In this study, we propose several simple techniques to replace a cross-lingual vocabulary with a compact, language-specific one.Namely, we address strategies for re-initializing the token embedding matrix after vocabulary specialization.We then provide a systematic experimental comparison of our techniques, in addition to the recently-proposed FOCUS method.We demonstrate that: 1) Embeddingreplacement techniques in the monolingual transfer literature are inadequate for adapting multilingual models.2) Replacing crosslingual vocabularies with smaller specialized ones provides an efficient method to improve performance in low-resource languages.3) Simple embedding re-initialization techniques based on script-wise sub-distributions rival techniques such as FOCUS, which rely on similarity scores obtained from an auxiliary model.