SciTechBaitRO: ClickBait Detection for Romanian Science and Technology News
Raluca-Andreea Gînga, Ana Sabina Uban · 2024
In this paper, we introduce a new annotated corpus of clickbait news in a low-resource language -Romanian, and a rarely covered domain -science and technology news: SciTech-BaitRO.It is one of the first and the largest corpus (almost 11,000 examples) of annotated clickbait texts for the Romanian language and the first one to focus on the sci-tech domain, to our knowledge.We evaluate the possibility of automatically detecting clickbait through a series of data analysis and machine learning experiments with varied features and models, including a range of linguistic features, classical machine learning (ML) models, deep learning and pre-trained models.We compare the performance of models using different kinds of features, and show that the best results are given by the BERT models, with results of up to 89% F1 score.We additionally evaluate the models in a cross-domain setting for news belonging to other categories (i.e.politics, sports, entertainment) and demonstrate their capacity to generalize by detecting clickbait news outside of domain with high F1-scores.