Stance Detection of Controversial Articles Using TF-IDF and BERT
Eka Parima Saragih, Anggraini Dyah Ayu Sekarlangit, Faqih Al Suman · Journal of Electrical Technology UMY · 2025
Online misinformation and polarized discussions require better methods for automatically detecting a text's stance. As digital content increases, identifying whether a news article supports, opposes, or is neutral towards its headline is crucial for fighting the spread of false information. This study presents a hybrid model designed for this task. We combine lexical features from Term Frequency-Inverse Document Frequency (TF-IDF), which captures word-level patterns, with contextual semantic information from a pretrained BERT model (bert-base-uncased). The features from both TF-IDF and BERT's [CLS] token were concatenated and used to train a logistic regression classifier. The model was trained and tested on a filtered version of the Fake News Challenge (FNC-1) dataset, with "unrelated" pairs removed to focus on more nuanced stance classification. The final evaluation of this model achieved 83% accuracy with a macro F1-score of 0.68. This model evaluates best in the Neutral stance (F1-score 0.91), but has some difficulty detecting the stance in the Oppositional class (with an F1-score 0.39). The results of this evaluation show that surface level lexical features combined with deep contextual understanding can improve the performance of stance detection.