Measuring Political Bias in LLMs using Fine-Tuning RoBERTa Model

V. Rajathi, A. Robert Singh, Karan Kumar · 2025

While the predominance of big language models (LLMs) on public conversation expands, their ability to identify and minimize political bias in their generated text is ever more vital. This research proposes a supervised approach which utilizes the RoBERTa model, with fine-tuning on a database of 37,554 labeled news stories marked manually for political bias (center, left, right). With a training-to-test ratio of 4:1, our model had an accuracy of 94.48% and macro F1-score of 94.49% after 4 epochs of training at a learning rate of 2e-5 and batch size of 8. The model shows robust performance in separating political orientations, with confidence scores of over ~98% for evident cases of bias. These findings validate that transformer-based models such as RoBERTa can effectively identify political bias in LLM-produced text, providing a basis for more transparent and responsible AI systems.

Read the paper · More papers on PaperTik