A comparative study on various pre-processing techniques and deep learning algorithms for text classification

P. Bhuvaneshwari, A. Nagaraja Rao · International Journal of Cloud Computing · 2022

Pre-processing is the primary technique employed in sentiment analysis and selecting the suitable techniques for the corresponding application can increase the classifier accuracy. It reduces the complexity innate in the raw data which makes the classifier to learn faster and precisely. Despite its importance, the pre-processing in polarity deduction has not attained much attention in the deep learning literature. So, in this paper, 13 popularly used pre-processing techniques are evaluated on three different domain online user review datasets. For evaluating the impact of each pre-processing technique, four deep neural networks are utilised, and they are auto-encoder, convolution neural network (CNN), long short-term memory (LSTM), and bidirectional LSTM (Bi-LSTM). The purpose of this paper is to identify the appropriate pre-processing techniques and the best classifier which achieves higher accuracy. Experimental results of this study show that using appropriate pre-processing techniques can significantly improve the classifier accuracy. Also, it is noted that Bi-LSTM model achieves higher accuracy rate than the remaining neural networks.

Read the paper · More papers on PaperTik