Prediction of toxicity-generating news using machine learning
Luís Leão Aguiar Braga da Cruz · Open Repository of the University of Porto (University of Porto) · 2020
The increased popularity of social networks in today's society, in combination with the facilitated acquisition of devices that provide access to those platforms, developed the conditions for the increase of users in these platforms.This growth sparked the attention of businesses such as news media outlets, which started enforcing their presence in social networks to provide users with news articles widening their audience.The use of social networks provides users with a costless and quick interaction, making interactions more impersonal and contributing to an increase in aggressive communication.Therefore, such interactions must be identifiable so that social networks can combat the use of aggressive communication harmful to its environment.Thereby, in this thesis, we investigate the problem of detecting toxicity-generating news.An initial objective was to make a review of the topic from a computer science perspective.We analysed the differences between toxicity and related concepts (hate, offensive and uncivil speech) and how they overlapped, complementing with our definition.Regarding past research, we could only find similar studies regarding the prediction of incendiary news, incivility and hate speech generating news.Despite the studies found, we concluded that, to our knowledge, there had been no studies on prediction of toxicity-generating news.With this conclusion, we decided to perform a review of studies on the area of news classification intending to review standard feature extraction techniques and machine learning algorithms used in the area.A second objective was the study of toxicity in comments and news present in our dataset.We used a dataset developed in the context of the Stop PropagHate project to predict hate speech generating news.The dataset contained news articles Twitter posts related to news media outlets from the USA, UK, Portugal and Brazil and respective user Twitter comments to those articles.We used the Perspective API to classify comments as toxic.From this classification, we concluded that the median toxicity in comments was 11.1%.With this metric, we considered news to be toxicity-generating if the mean toxicity of the comments towards that article was equal or higher than the median toxicity.As a final objective, we intended to predict news as toxicity-generating and understand which features contributed for the prediction.For this goal, we extracted meta-data features and newscontent features.We conducted experiments with training, validation and test phases.We performed several feature combination experiments and concluded that our best model represented a combination of meta-data features and news content features, reaching an F1 score of 0.74.Furthermore, analysing the feature correlation, we concluded that the model's performance was not resultant of a subset of features, but rather the combination of all.Regarding the features that most contributed to the classification of toxicity-generating news, we concluded that features such as the number of comments a news originated, and title keywords, were the most significant contributors to the classification.Moreover, we concluded that articles which titles contained keywords relative to highly debated social topics such as "racist", "gay" and "nazi" in conjunction to political entities such as "Trump" contributed to the identification of toxicity-generating news.All the mentioned points and objectives were met.We successfully developed a model reasonably capable of classifying toxicity-generating news, and understand factors behind the phenomena.i Luís Braga da Cruz v vi "Success is not final, failure is not fatal: it is the courage to continue that counts."