ChatGPT vs. Human Journalists: Analyzing News Summaries Through BERTScore and Moderation Standards

Hui-Sang Kim, Ji-Won Kang, Sun‐Yong Choi · Electronics · 2025

Recent advances in natural language processing (NLP) have enabled the development of powerful language models such as Generative Pre-trained Transformers (GPTs). This study evaluates the performance of ChatGPT in generating news summaries by comparing them with summaries written by professional journalists at The New York Times. Using BERTScore as the primary metric, we assessed the semantic similarity between ChatGPT-generated and human-authored summaries. We further employed OpenAI’s moderation API to examine the extent to which each set of summaries contained potentially biased, inflammatory, or violent language. The results indicate that ChatGPT-generated summaries exhibit a high degree of contextual alignment with human-written summaries, achieving a BERTScore F1-score above 0.87. Moreover, ChatGPT outputs consistently omit language flagged as problematic by moderation algorithms, producing summaries that are less likely to include harmful or polarizing content—a feature we define as moderation-friendly summarization. These findings suggest that ChatGPT can serve as a valuable tool for automated news summarization, offering content that is both contextually accurate and aligned with content moderation standards, thereby supporting more objective and responsible news dissemination.

Read the paper · More papers on PaperTik