Evaluation of Reliability Criteria for News Publishers with Large Language Models
Manuel Pratelli, John Bianchi, Fabio Pinelli, Marinella Petrocchi · 2025
In this study, we investigate the use of a large language model to assist in the evaluation of the reliability of the vast number of existing online news publishers, addressing the impracticality of relying solely on human expert annotators for this task.In the context of the Italian news media market, we first task the model with evaluating expert-designed reliability criteria using a representative sample of news articles.We then compare the model's answers with those of human experts.The dataset consists of 352 news articles annotated by three human experts and the LLM.Examining 6,081 annotations over six criteria, we observe good agreement between LLM and human annotators in three evaluated criteria, including the critical ability to detect instances where a text negatively targets an entity or individual.For two additional criteria, such as the detection of sensational language and the recognition of bias in news content, LLMs generate fair annotations, albeit with certain trade-offs.Furthermore, we show that the LLM is able to help resolve disagreements among human experts, especially in tasks such as identifying cases of negative targeting.