Sentiment Analysis in Call Center Conversations Using Large Language Models and BERT
Yusuf Sali, Sıtkı Can Toraman · 2025
Although sentiment analysis (SA) has been widely studied in various domains such as social media, e-commerce, and movie and book reviews, research on SA in dialogues remains highly limited. In Task Oriented Dialogue Systems (TODS), such as call center conversations, SA is a highly valuable task that can provide significant automation benefits. In this study, a total of 12 different models, including Large Language Models (LLMs) such as GPT-4o and o3-mini, as well as BERT and RoBERTa, were used for SA on TODS data. Additionally, to examine the impact of training datasets in domains where publicly available datasets do not exist, the Winvoker dataset and a synthetic dataset generated using GPT-4o were compared with an indomain dataset annotated by experts. The findings indicate that when LLMs are used in a few-shot instruction-based manner without fine-tuning, their performance falls significantly below that of BERT models.