Distill, Adapt, Distill: Training Small, In-Domain Models for Neural Machine Translation

Mitchell A. Gordon, Kevin Duh · 2020

We explore best practices for training small, memory efficient machine translation models with sequence-level knowledge distillation in the domain adaptation setting.While both domain adaptation and knowledge distillation are widely-used, their interaction remains little understood.Our large-scale empirical results in machine translation (on three language pairs with three domains each) suggest distilling twice for best performance: once using general-domain data and again using indomain data with an adapted teacher.The code for these experiments can be found here.1

Read the paper · More papers on PaperTik