Transfer Learning for Digital Heritage Collections: Comparing Neural Machine Translation at the Subword-level and Character-level

Nikolay Banar, Karine Lasaracina, Walter M. P. Daelemans, Mike Kestemont · 2020

Transfer learning via pre-training has become an important strategy for the efficient application of NLP methods in domains where only limited training data is available.This paper reports on a focused case study in which we apply transfer learning in the context of neural machine translation (French-Dutch) for cultural heritage metadata (i.e.titles of artistic works).Nowadays, neural machine translation (NMT) is commonly applied at the subword level using byte-pair encoding (BPE), because word-level models struggle with rare and out-of-vocabulary words.Because unseen vocabulary is a significant issue in domain adaptation, BPE seems a better fit for transfer learning across text varieties.We discuss an experiment in which we compare a subword-level to a character-level NMT approach.We pre-trained models on a large, generic corpus and fine-tuned them in a two-stage process: first, on a domain-specific dataset extracted from Wikipedia, and then on our metadata.While our experiments show comparable performance for character-level and BPEbased models on the general dataset, we demonstrate that the character-level approach nevertheless yields major downstream performance gains during the subsequent stages of fine-tuning.We therefore conclude that character-level translation can be beneficial compared to the popular subword-level approach in the cultural heritage domain.

Read the paper · More papers on PaperTik