A Novel Framework for Text Preprocessing using NLP Approaches and Classification using Random Forest Grid Search Technique for Sentiment Analysis
SHRIVASH Brajesh Kumar, VERMA Dinesh Kumar, PANDEY Prateek · ECONOMIC COMPUTATION AND ECONOMIC CYBERNETICS STUDIES AND RESEARCH · 2025
Text preprocessing is a process to organize and prepare raw text for machine learning (ML) and deep learning (DL) models.This process is most widely used in Sentiment Analysis (SA).To process raw text data, Natural Language Processing (NLP) plays a vital role in data preprocessing by cleaning text, removing punctuations and stopwords, stemming and lemmatization.The classification of text data is a common use of NLP.An effective data preprocessing approach also identifies text features efficiently.The most essential component of text classification is feature extraction from raw text data to input into ML and DL models.Feature extraction prepares training data sets effectively for ML and DL models to find good results.In this paper, we proposed a novel framework for text preprocessing using NLP approaches, and feature extraction, and expanded it for the Random Forest ML classifier.To improve the performance of the random forest classifier, we examined the grid search technique with estimator Random Forest Classifier and measured the performance of the model.The accuracy measure for the random forest model was 93%, while the accuracy measure using the grid search technique was 94%, which shows that the grid search technique enhances the model performance.