Social Media Text Classification For Hate Speech Detection Using Different Feature Selection Techniques

Priyanshu Jadon, Deepshikha Bhatia, Durgesh Kumar Mishra · 2024

Social media content is one of the unregulated data and there is not any monitory policy is available for controlling the contents. However, a number of techniques are developed for monitoring the text data in social media but due to lake of effective feature extraction techniques there are very less accuracy has been observed in this context. In this paper we have tried to measure the impact of different feature-extraction methods over the deep learning model. In this context we have implemented five different feature selection models namely bi-gram based features, PoS based features, word count vectorizer, TF-IDF based Features and word embedding technique. These features are trained using a deep learning neural network and their performances of classification have been measured. The experimentation of hate speech detection dataset and the extensive experimental analysis we have found the following facts (1) the word count vectorizer and TF-IDF based features are providing higher training and validation accuracy (2) the variation in feature vector dimension can increase or decrease the classification accuracy (3) the higher dimensional feature vectors can increase resource consumption in terms of time and memory (4) in order to classify the text on social media the words may have grate importance. Based on this investigation we have planned to design a new feature descriptor that will improve the classification performance of small social media text.

Read the paper · More papers on PaperTik