Research on Predicting Public Opinion Event Heat Levels Based on Large Language Models
Yi Ren, Tianyi Zhang, Weibin Li, DuoMu Zhou, Chenhao Qin, FangCheng Dong · 2024
With the rapid advancement of large language models, several models, such as GPT-4o, have demonstrated extraordinary capabilities, surpassing human performance in various language tasks. This study proposes a novel approach using large language models for public opinion event heat level prediction. We preprocessed and classified 62,836 Chinese hot event data collected between July 2022 and December 2023. Based on each event’ s online dissemination heat index, we used the MiniBatchKMeans algorithm to cluster these events into four heat levels (from low heat to very high heat). We then randomly selected 250 events from each level, totaling 1,000 events, to form the evaluation dataset. In testing, we assessed the accuracy of various language models in predicting event heat levels in two scenarios: without reference cases and with similar case references. The results showed that GPT-4o and DeepseekV2 performed best in the referenced case scenario, achieving prediction accuracies of 41.4% and 41.5%, respectively. Although overall prediction accuracy remains relatively low, for low-heat (Level 1) events, GPT-4o and DeepseekV2 achieved accuracies of 73.6% and 70.4%, respectively. Additionally, prediction accuracy showed a downward trend from Level 1 to Level 4, corresponding to the uneven data distribution across heat levels. This suggests that with more robust datasets, public opinion event heat level prediction using large language models holds significant research potential for the future.