Leveraging emotion and word-based features for antisocial behavior detection in user-generated content
Myriam Munezero · UEF eRepo (University of Eastern Finland) · 2017
Online platforms increasingly provide opportunities to publish user-generated content that expresses emotions, thoughts, and intentions.Some of this content may include emotions, thoughts, and intentions related to antisocial behavior.With the vast amount of user content generated each day, it has become overwhelming and impossible to manually monitor and detect incidents of antisocial behavior.While antisocial behavior, which is harmful to society, has often been studied from educational, social, and psychological points of view, little research has been conducted on computational linguistics and natural language processing regarding antisocial behavior.To address this gap, this work investigated whether recent advances in natural language processing methods and tools can be used to automatically detect potential instances of antisocial behavior.This dissertation presents a framework that can be used for the automatic detection of antisocial behavior in text.The framework is based on the emotion and language theories, which provide a comprehensive understanding of antisocial behavior and explain how antisocial behavior, word usage, writing styles, and emotions are connected and are represented in text.The framework leverages word-based and emotion features for the detection of antisocial behavior in user-generated content.Using the features, supervised machine learning models were developed for the automatic detection of antisocial behavior.The effectiveness of the developed models was evaluated based on a set of corpora, one of which -the antisocial behavior corpus -was created by the author of this research and her colleagues.Several of the developed models showed high accuracies of over 90% in their detection of antisocial behavior in text.In addition, the distinguishing features of antisocial behavior texts, such as a high use of swearing, insults, and negative emotions, including anger, were identified and analyzed, thus allowing for an improved understanding of antisocial behavior in user-generated content.In sum, this study has demonstrated the potential of utilizing natural language processing techniques for antisocial behavior detection.Continued research on the relationships between natural language use and public security concerns as well as multidisciplinary efforts to develop models that can accurately predict harmful behavior are required to extend this research.