Improving BERT With Self-Supervised Attention
Yiren Chen, Xiaoyu Kou, Jiangang Bai, Yunhai Tong · IEEE Access · 2021
One of the most popular paradigms of applying large pre-trained NLP models such as BERT is to fine-tune it on a smaller dataset. However, one challenge remains as the fine-tuned model often overfits on smaller datasets. This problem manifests as irrelevant or misleading words in the sentences, which are obvious to humans, but can substantially degrade the performance of the fine-tuned BERT models. In this paper, we propose a novel technique, called Self-Supervised Attention (SSA) to help facilitate this generalization challenge. Specifically, SSA automatically generates weak, token-level attention labels iteratively by probing the fine-tuned model from the previous iteration.We investigate two different ways of integrating SSA into BERT and propose a hybrid approach to combine their benefits. Empirically, through a variety of public datasets, we illustrate significant performance improvement using our SSA-enhanced BERT model.