Sentiment Bias and Security Analysis in Training Datasets of Large Language Models
Keito Inoshita, Xiaokang Zhou · 2024
The development of large language models (LLMs) has led to improved predictive performance and efficiency in various fields. However, the datasets used for training LLMs contain country-specific sentiment biases, which could pose serious issues in the field of security. This study aims to analyze the country-specific sentiment biases in the datasets used for training LLMs and evaluate their impact. In this study, we analyzed the sentiment tendencies associated with country names using VADER across four major large-scale training datasets: C4, RedPajama, OSCAR, and RefinedWeb. Specifically, we calculated the frequency of appearance and sentiment tendencies for each country, and further assessed the sentiment tendencies in bilateral relationships. The results revealed that the USA had a significantly high frequency of appearances and exhibited greater robustness compared to other countries. In contrast, negative sentiments towards Russia and Iran were prominently emphasized. Additionally, notable sentiment biases were confirmed in specific combinations of bilateral relationships. These findings indicate the presence of substantial biases in the datasets used for training LLMs. Particularly in the field of security, there is a risk of misunderstandings and conflicts arising in international relations. To address this, it is necessary to create datasets that evaluate LLMs’ perceptions of national conflicts and establish bias-removal standards in collaboration with international organizations. This will facilitate the development of more fair and reliable AI systems.