MA-BERT: Learning Representation by Incorporating Multi-Attribute Knowledge in Transformers

You Zhang, Jin Wang, Liang-Chih Yu, Xuejie Zhang · 2021

Incorporating attribute information such as user and product features into deep neural networks has been shown to be useful in sentiment analysis.Previous works typically accomplished this in two ways: concatenating multiple attributes to word/text representation or treating them as a bias to adjust attention distribution.To leverage the advantages of both methods, this paper proposes a multi-attribute BERT (MA-BERT) to incorporate external attribute knowledge.The proposed method has two advantages.First, it applies multi-attribute transformer (MA-Transformer) encoders to incorporate multiple attributes into both input representation and attention distribution.Second, the MA-Transformer is implemented as a universal layer and stacked on a BERT-based model such that it can be initialized from a pre-trained checkpoint and fine-tuned for the downstream applications without extra pretraining costs.Experiments on three benchmark datasets show that the proposed method outperformed pre-trained BERT models and other methods incorporating external attribute knowledge.

Read the paper · More papers on PaperTik