Novel Approach for Generating Hybrid Features Set to Effectively Identify Hate Speech
Shruthi P, Anil Kumar K M · INTELIGENCIA ARTIFICIAL · 2020
Automating hate speech or inappropriate text detection in social media and other internet platforms isgaining a lot of interest and becoming a valuable research topic for both industry and academia in recent years. Itis more important for applications to identify the disruptive contents, understand sentiment analysis, identify cyberbullying, detect flames, threats, hatred towards people or particular communities or groups etc. Text classificationis a very challenging task due to the nature and complexities with languages, especially its context, micro words,emojis, typo error and sarcasm present in the text. In this paper, we have proposed a model with a novel approachfor generating hybrid features for an effective feature representation to classify hate speech. We have combinedfeatures learned from deep learning methods with the semantic features like word n-grams and tweets specificsyntactic features to form hybrid feature sets. We have also improvised preprocessing steps to reduce the numberof missing embeddings to increase the vocabulary for efficient feature learning. We have experimented with thevarious neural networks for feature learning and machine learning models with hybrid features for classification.Our work delivers hybrid features and appropriate preprocessing techniques for an efficient classification of thestandard dataset of 16k annotated hate speech tweets. The combination of Long Short Term Memory (LSTM)trained on Random Embeddings for deep learning features extraction and Logistic Regression (LR) as a classifierwith the hybrid features is found to be the best model and it outperforms the state of the art reported in theliterature.