Fight for 4230 at CheckThat! 2021: Domain-Specific Preprocessing and Pretrained Model for Ranking Claims by Check-Worthiness
Xinrui Zhou, Bohuai Wu, Pascale Fung · Rare & Special e-Zone (The Hong Kong University of Science and Technology) · 2021
The widespread dissemination of false news on social media has brought negative effects to society. In this paper, we describe a model submitted to the CLEF-2021 CheckThat! Task 1 - English to estimate the check-worthiness of tweets and political debates/speeches. Our official submission was ranked 2nd in subtask 1A with a MAP score of 0.195 and ranked 1st in subtask 1B with a MAP score of 0.402. The main challenges of the task 1 are the imbalanced data and the not standard texts of tweets. We did thorough data preprocessing and mainly focused on combining different pretrained models with a dropout layer and a dense linear layer. We explored and experimented with many combinations of different data preprocessing techniques and augmentation methods. We also tried extracting additional features from metadata and ensembling the best-performance models to further improve. We have developed a preprocessing procedure for tweets, and our experiments show that domain-specific preprocessing and pretrained models can significantly improve the performance. Finally, we submitted the result produced by the BERTweet model with extra dropout layer and classifier layer with preprocessed data for subtask 1A and RoBERTa model fine-tuned on tweets_hate_speech_detection dataset with extra dropout layer and classifier layer for subtask 1B. © 2021 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).