Multi-Label Classification of CS Papers Using Natural Language Processing Models
Zheng Zhang, Amit Kumar Das, Mostafa Rahgouy, Yida Bao, Sanjeev Baskiyar · 2023
Computer science (CS) is a rapidly evolving field of knowledge with a significant volume and significance of scientific papers. Consequently, meticulous management and categorization of CS papers are crucial to ensure efficient access and retrieval of relevant information. This study aims to evaluate the effectiveness of eight traditional machine learning models, as well as Recurrent Neural Network (RNN) models, such as LSTM and BiLSTM, and state-of-the-art pre-trained self-attention-based natural language processing (NLP) models, including BERT, XLNet, RoBERTa, and DistilBERT, for the task of multi-label classification of arXiv CS paper classification. The results demonstrate that pre-trained self-attention-based models consistently outperform the other models regarding classification performance. Moreover, self-attention-based models exhibited superior performance and achieved a new state-of-the-art result. To the best of our knowledge, our paper represents the first endeavor to evaluate machine learning models, specifically on computer science-related documents, as part of a multi-label classification task.