Better Pre-Training by Reducing Representation Confusion

Haojie Zhang, Mingfei Liang, Ruobing Xie, Zhenlong Sun, Bo Zhang, Leyu Lin · 2023

In this work, we revisit the Transformer-based pre-trained language models and identify two different types of information confusion in position encoding and model representations, re-* Equal contribution. 1 "attention weights" mainly refer to the dot product between Key and Query in the self-attention module.2 MTH is the abbreviation of our proposed MLM with Token Cosine Differentiation (TCD) and Head Cosine Differentiation (HCD) pre-training task.TCD and HCD are described in detail in sec.1(2) and sec.3.2.

Read the paper · More papers on PaperTik