LLM-based Adversarial Dataset Augmentation for Automatic Media Bias Detection

Martin Paul Wessel · 2025

This study presents BiasAdapt, a novel data augmentation strategy designed to enhance the robustness of automatic media bias detection models.Leveraging the BABE dataset, Bi-asAdapt uses a generative language model to identify bias-indicative keywords and replace them with alternatives from opposing categories, thus creating adversarial examples that preserve the original bias labels.The contributions of this work are twofold: it proposes a scalable method for augmenting bias datasets with adversarial examples while preserving labels, and it publicly releases an augmented adversarial media bias dataset.Training on Bi-asAdapt reduces the reliance on spurious cues in four of the six evaluated media bias categories.

Read the paper · More papers on PaperTik