PatchBERT: Just-in-Time, Out-of-Vocabulary Patching

Sangwhan Moon, Naoaki Okazaki · 2020

Large scale pre-trained language models have shown groundbreaking performance improvements for transfer learning in the domain of natural language processing.In our paper, we study a pre-trained multilingual BERT model and analyze the OOV rate on downstream tasks, how it introduces information loss, and as a side-effect, obstructs the potential of the underlying model.We then propose multiple approaches for mitigation and demonstrate that it improves performance with the same parameter count when combined with finetuning.

Read the paper · More papers on PaperTik