Do LLMs Understand Dialogues? A Case Study on Dialogue Acts
Ayesha Qamar, Jonathan Tong, Ruihong Huang · 2025
Recent advancements in NLP, largely driven by Large Language Models (LLMs), have significantly improved performance on an array of tasks.However, Dialogue Act (DA) classification remains challenging, particularly in the fine-grained 50-class, multiparty setting.This paper investigates the root causes of LLMs' poor performance in DA classification through a linguistically motivated analysis.We identify three key pre-tasks essential for accurate DA prediction: Turn Management, Communicative Function Identification, and Dialogue Structure Prediction.Our experiments reveal that LLMs struggle with these fundamental tasks, often failing to outperform simple rulebased baselines.Additionally, we establish a strong empirical correlation between errors in these pre-tasks and DA classification failures.A human study further highlights the significant gap between LLM and human-level dialogue understanding.These findings indicate that LLMs' shortcomings in dialogue comprehension hinder their ability to accurately predict DAs, highlighting the need for improved dialogue-aware training approaches. InstructionYou are an intelligent annotator capable of classifying the intention behind each speaker's utterance.You will be provided with a list of possible Dialogue Acts and their definitions.You will be given an utterance surrounded by '#'.Your task is to predict the correct label for that utterance.You will also be given a snapshot of the conversation to provide context for your prediction.Return the answer in the format: 'label:predicted label'.The Dialogue Acts and their definitions are as follows: Statement: General statements.Accept: A short utterance indicating acceptance of a previous speaker's statement.Disruption: Indecipherable or disrupted speech.Defending/Explanation: The speaker defends their opinion or provides an explanation.