Test of NLP models with various news datasets

Hyunwoo Ko, Jun-Kwon Hwangbo · The Journal of the Korean Institute of Information and Communication Engineering · 2023

Natural Language Processing is one of the fields that attracts a lot of attention in deep learning, and with the introduction of transformer-based GPT[1] and BERT[2], it is showing tremendous performance improvement. In this paper, we compared and analyzed the performance of word embedding, neural network, and pre-trained language model, dependent othe model and data type of the news. ISOT, Kaggle, and Politifact datasets were used for fake news dataset as a result, BERT showed best performance in this study, however in Politifact Dataset, it showed relatively poor performance. We analyzed the structure of dataset and from the model perspectives to find out the reason why the performance differences were occurred.

Read the paper · More papers on PaperTik