jhan014 at SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media

Jiahui Han, Shengtan Wu, Xinyu Liu · 2019

In this paper, the team jhan014 presents two methods to identify and categorize the offensive language in Twitter.In the first method, we develop a deep neural network consisting of bidirectional recurrent layers with Gated Recurrent Unit (GRU) cells and fully connected layers.In the second method, we establish a probabilistic model, modified sentence offensiveness calculation (MSOC) to evaluate the sentence offensiveness level and target level according to different sub-tasks.Based on task results, We evaluate the performance of each method based on F1 score and analyze the advantages and disadvantages of these two methods with the type I error and type II error.In conclusion, deep neural network behaves well in all subtasks but has more type I error and fails to categorize subclasses with minor data or less character, while MSOC model does better in target categorizing but has more type II error in offensive identifying.

Read the paper · More papers on PaperTik