Do not Mask Randomly: Effective Domain-adaptive Pre-training by Masking In-domain Keywords
Shahriar Golchin, Mihai Surdeanu, Nazgol Tavabi, Ata M. Kiapour · 2023
We propose a novel task-agnostic in-domain pre-training method that sits between generic pre-training and fine-tuning.Our approach selectively masks in-domain keywords, i.e., words that provide a compact representation of the target domain.We identify such keywords using KeyBERT (Grootendorst, 2020).We evaluate our approach using six different settings: three datasets combined with two distinct pretrained language models (PLMs).Our results reveal that the fine-tuned PLMs adapted using our in-domain pre-training strategy outperform PLMs that used in-domain pre-training with random masking as well as those that followed the common pre-train-then-fine-tune paradigm.Further, the overhead of identifying in-domain keywords is reasonable, e.g., 7-15% of the pretraining time (for two epochs) for BERT Large (Devlin et al., 2019). 1