Feature Extraction from Text using Grasshopper optimization algorithm for identifying hate speech

Anu Dhankhar, Amresh Prakash, Sapna Juneja · 2023

The optimization algorithms have ability to eliminate features that are extraneous and redundant, shrink the vector space, manage computation time, and enhance performance on problems requiring more accurate classification. Feature selection and feature extraction have always been crucial, especially in text categorization. Using optimization algorithms, engineering methods are used in these can be improved still further. Using five machine learning (ML) models—Decision Tree (DT), Naive Bayes (NB), Random Forest (RF) and Multilayer Perceptron (MLP) and Deep Neural Network (DNN)—this paper implements one such optimization algorithm, Grasshopper optimization (GOA), using various feature extraction and selection methods on textual datasets. The machine learning model performs more effectively due to the suggested feature selection and feature extraction strategies. In order to identify hate speech, this study takes into consideration text-based datasets. The positive, negative, and neutral sentiments of tweets are extracted from the Twitter API to create the text collection. On the application of GAO, Combining Random Forest with the method of feature extraction TF-IDF shows a greatest accuracy improvement of 24% as compared with without applied any optimization algorithm and 13% accuracy increased as compared with another optimization algorithm i.e., ACO. GOA provides an accuracy of up to 93%, which is superior to ACO optimization algorithm's 80% accuracy when used with the Random Forest model.

Read the paper · More papers on PaperTik