Comparing BERT and XLNet from the Perspective of Computational Characteristics

Hailong Li, Jaewan Choi, Sunjung Lee, Jung Ho Ahn · 2020 International Conference on Electronics, Information, and Communication (ICEIC) · 2020

Exploiting attention mechanism, Transformer provides superior performance compared to traditional CNN and RNN models on various NLP (Natural Language Processing) tasks. BERT and XLNet are two popular models utilizing Transformer. In this paper, we compare the computational characteristics of the inference of BERT and XLNet using MPRC (Microsoft Research Paraphrase Corpus), one of the popular language understanding benchmarks. Through evaluation, we observe that the both models exhibit similar computational characteristics except the target-position-aware representation and relative position encoding features of XLNet, leading to a better benchmark score at the cost of$\mathit{1.2}\times$arithmetic operations and$\mathit{1.5}\times$execution time on a modern CPU.

Read the paper · More papers on PaperTik