Domain adapted machine translation: What does catastrophic forgetting forget and why?
Danielle Saunders, Steve DeNeefe · 2024
Neural Machine Translation (NMT) models can be specialized by domain adaptation, often involving fine-tuning on a dataset of interest.This process risks catastrophic forgetting: rapid loss of generic translation quality.Forgetting has been widely observed, with many mitigation methods proposed.However, the causes of forgetting and the relationship between forgetting and adaptation data are under-explored.This paper takes a novel approach to understanding catastrophic forgetting during NMT adaptation by investigating the impact of the data.We provide a first investigation of what is forgotten, and why.We examine the relationship between forgetting and the in-domain data, and show that the amount and type of forgetting is linked to that data's target vocabulary coverage.Our findings pave the way toward better informed NMT domain adaptation.