A Parameter-Adaptive Convolution Neural Network for Capturing the Context-Specific Information in Natural Language Understanding
Rongcheng Duan, Xiaodian Yang, Qingyu Wang, Yang Zhao · 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT) · 2021
Natural Language Understanding (NLU) aims to make sense of language by enabling computers to comprehend text in semantic level, which is a fundamental but challenging task in natural language processing. Recently, BERT has achieved state-of-the-art performances in NLU utilizing pretraining and fine-tuning techniques to capture the task-specific information without any task-specific structure. However, the existing models still cannot capture context-specific information, leading to poor performance when dealing with the pervasive ambiguity of language caused by the polysemous words. In this paper, we propose a Parameter-Adaptive Convolution Neural Network (PACNN) to capture the context-specific information which can deal with the polysemous word better. Instead of the convolution layers in existing models, the parameters of convolutional filters in PACNN, generated by a deconvolution (e.g. convolution transpose) neural network, are adaptable according to the input sentences, which can filter the local information via the global information and capture the context-specific meaning of the polysemous words. We empirically demonstrate the efficiency of the proposed PACNN by performing a series of experiments on the General Language Understanding Evaluation (GLUE) benchmark, a collection of popular datasets on different tasks, and the PACNN significantly outperforms all the baselines. Besides, our model can be appended to BERT to further improve its performances. As BERT and the proposed PACNN capture information from various aspects, the proposed BERT+PACNN achieves the best performances compared with BERT and other baselines. Furthermore, we visualize the task-specific information and context-specific information captured by BERT and the PACNN, separately, using Singular Value Decomposition (SVD) to demonstrate the efficiencies of the two models further.