Pre-Training Transformers as Energy-Based Cloze Models

Kevin B. Clark, Minh-Thang Luong, Quoc Viet Le, Christopher D. Manning · 2020

We introduce Electric, an energy-based cloze model for representation learning over text.Like BERT, it is a conditional generative model of tokens given their contexts.However, Electric does not use masking or output a full distribution over tokens that could occur in a context.Instead, it assigns a scalar energy score to each input token indicating how likely it is given its context.We train Electric using an algorithm based on noise-contrastive estimation and elucidate how this learning objective is closely related to the recently proposed ELECTRA pre-training method.Electric performs well when transferred to downstream tasks and is particularly effective at producing likelihood scores for text: it reranks speech recognition n-best lists better than language models and much faster than masked language models.Furthermore, it offers a clearer and more principled view of what ELECTRA learns during pre-training.

Read the paper · More papers on PaperTik