Empowering Topic Classification Accuracy by Integrating Transformer-Based with Machine Learning and Deep Learning Techniques

Rana Reda Waly, Mosatafa Mohamed Saeed, Mena Hany, Abdelrahman Ezzeldin Nagib, Wael Hassan Gomaa · 2023

Identifying the topic of a piece of text is a crucial responsibility in natural language processing and the analysis of written language. This study investigates the performance of two approaches - machine learning and deep learning - for topic classification. In the machine learning approach, we tested ten different classifiers after applying various pre-processing techniques to our dataset and used T5 X-Large as the pre-trained embedding model. The best result was achieved using the Support Vector Machine (SVM) classifier with minimal pre-processing, i.e., only removing special characters. This method obtained an accuracy of 89%. While in the deep learning approach, we experimented with different pre-trained models, pre-processing techniques, and input layers for our neural network architecture. The optimal configuration was found using a Recurrent Neural Network (RNN) as the input layer, applying lowercase as the only pre-processing technique, and employing T5 X-Large as the embedding model. This configuration achieved an accuracy score of 89.91%. The results demonstrate that both approaches have their merits, but the deep learning approach offers a slight improvement in classification performance. This research provides valuable insights for researchers and practitioners working on topic classification tasks in natural language processing and text analysis.

Read the paper · More papers on PaperTik