Multi-Class Text Classification using Machine Learning & Deep Learning

Arul Keswani, Tarun Kumar Jain, Bibek Sharma · 2023

Text classification indeed holds a central position in the field of natural language processing (NLP) and has a wide range of applications across diverse domains. In this paper, we investigate the effectiveness of BERT & DistilBERT embeddings in combination with long short-term memory (LSTM), convolutional neural networks (CNN), and bi-directional LSTM (bi-LSTM) architectures for text classification on the widely recognized 20 Newsgroups dataset. We propose an approach that utilizes BERT & DistilBERT embeddings to capture contextual information and feed it into the CNN, LSTM, and bi-LSTM models. This combination enables the models to effectively learn hierarchical representations and capture both local and global dependencies in the text data. Experimental results demonstrate that our proposed CNN, LSTM, and bi-LSTM models outperform Siamese networks, which have previously been explored for text classification on the 20 Newsgroups dataset. Our models consistently demonstrate superior accuracy and deliver enhanced performance in terms of precision, recall, and F1-score across diverse classes. Our proposed models not only outperform Siamese networks but also provide insights into the importance of leveraging contextual embeddings for more accurate and robust text classification. A maximum accuracy of 92.94% was achieved when using distil-BERT with bi-LSTM and a minimum of 87.76% when using LSTM.

Read the paper · More papers on PaperTik