Context Analysis for Pre-trained Masked Language Models
Yi-An Lai, Garima Lalwani, Yi Zhang · 2020
Pre-trained language models that learn contextualized word representations from a large unannotated corpus have become a standard component for many state-of-the-art NLP systems.Despite their successful applications in various downstream NLP tasks, the extent of contextual impact on the word representation has not been explored.In this paper, we present a detailed analysis of contextual impact in Transformer-and BiLSTM-based masked language models.We follow two different approaches to evaluate the impact of context: a masking based approach that is architecture agnostic, and a gradient based approach that requires back-propagation through networks.The findings suggest significant differences on the contextual impact between the two model architectures.Through further breakdown of analysis by syntactic categories, we find the contextual impact in Transformer-based MLM aligns well with linguistic intuition.We further explore the Transformer attention pruning based on our findings in contextual analysis.