Improving Natural Language Understanding by Reverse Mapping Bytepair Encoding
Chaodong Tong, Huailiang Peng, Qiong Dai, Lei Jiang, Jianghua Huang · 2019
Recently, language models (LMs) or language representation models are widely used in natural language understanding (NLU) tasks.However, these LMs are usually trained on large unlabeled text corpora, while the finetuning process simply takes words or wordpieces as model input.Because of the differences between language model and NLU task objectives, the problem of lack of concern on some key words exists.Thus in this paper, we propose a method called reverse mapping bytepair encoding, which maps named-entity information and other word-level linguistic features back to subwords during the encoding procedure of bytepair encoding (BPE).We employ this method to the Generative Pre-trained Transformer (OpenAI GPT) (Radford et al., 2018) by adding a weighted linear layer after the embedding layer.We also propose a new model architecture named as the multi-channel separate transformer to evaluate the effectiveness of the newly introduced information by employing a training process without parameter-sharing.Experiments on Story Cloze, RTE, SciTail and SST-2 datasets demonstrate the effectiveness of our approach.Compared with the original results in GPT, our approach gains 1.58% absolute increase on Stories Cloze, 6.4% on RTE, 0.69% on SciTail and 0.8% on SST-2.