LiBERTy: A Novel Model for Natural Language Understanding

Onkar Susladkar, Gayatri Deshmukh, Sparsh Mittal, Sai Chandra Teja R, Rekha Singhal · 2023

Recent advances in pre-trained neural language models have substantially enhanced the performance of numerous natural language processing (NLP) tasks. However, some existing models require pretraining on a large dataset. Moreover, on using a deep network with sequentially connected transformer blocks, there is a data loss across these blocks. To overcome these challenges, we propose LiBERTy, a novel network for natural language understanding. LiBERTy uses a novel TransLSTM module which takes the representations from the BERT block as input and feeds it to LSTM which functions as a pooling layer. The use of LSTM as a pooler helps the model sequentially encode the feature map into hidden states and understand semantic interrelations. The output of the TransLSTM module is fed to a classifier, which uses multiple 1D-CONV blocks, a 1D adaptive average pooling layer and a “fully-connected” (FC) layer and then, ArcFace loss. ArcFace loss helps in achieving inter-class separability and intra-class compactness. Our proposed strategies increase the efficiency of model pre-training and the performance of both natural language understanding (NLU) and downstream tasks. We showcase the efficacy of LiBERTy by applying it for three tasks: (1) disaster tweet classification on the HumAID dataset, (2) fine-grained emotion analysis on the GoEmotions dataset and (3) named entity recognition on TASTEset dataset. On all these datasets, LiBERTy provides comparable or superior F1-score compared to state-of-art networks. The source code is available at https://github.com/CandleLabAI/Liberty-NLP-model-2023.

Read the paper · More papers on PaperTik