Beyond Recognising Entailment: Formalising Natural Language Inference from an Argumentative Perspective

Ameer Saadat-Yazdi, Nadin Kökciyan · 2024

In argumentation theory, argument schemes are a characterisation of stereotypical patterns of inference.There has been little work done to develop computational approaches to identify these schemes in natural language.Moreover, advancements in recognizing textual entailment lack a standardized definition of inference, which makes it challenging to compare methods trained on different datasets and rely on the generalisability of their results.In this work, we propose a rigorous approach to align entailment recognition with argumentation theory.Wagemans' Periodic Table of Arguments (PTA), a taxonomy of argument schemes, provides the appropriate framework to unify these two fields.To operationalise the theoretical model, we introduce a tool to assist humans in annotating arguments according to the PTA.Beyond providing insights into non-expert annotator training, we present Kialo-PTA24, the first multi-topic dataset for the PTA.Finally, we benchmark the performance of pre-trained language models on various aspects of argument analysis.Our experiments show that the task of argument canonicalisation poses a significant challenge for state-of-the-art models, suggesting an inability to represent argumentative reasoning and a direction for future investigation.1. We conduct an annotation study to rephrase natural language arguments into structured templates and provide insights into how to train non-expert annotators to perform this analysis.2. We introduce ArgNotator, a tool that assists humans in annotating arguments according to the PTA.3. We construct Kialo-PTA24 -the first multi-topic dataset of argument types annotated according to the PTA.4. We compare the performance of state-of-theart models for two annotation subtasks.For the substance classification task, we benchmark the performance of a number of BERT-based models.For the argument canonicalisation task, we evaluate the performance of two large language models (FLAN-T5, LLAMA2) in both pre-trained and few-shot settings.The dataset, experimental setup, annotation tool and training materials can all be found on GitLab. 1

Read the paper · More papers on PaperTik