An Automatic Method for Standartizing Argumentative Annotations across Annotators

Ivan Sergeevich Pimenov, Natalia V. Salomatina · 2024

The prevalence of machine learning methods in addressing natural language processing tasks, particularly the problem of automatic extraction of argumentation structures, entails a demand for developing sufficient-sized text corpora with reliable argumentation annotation. Annotation reliability implies consistency in annotating analyzed entities, especially in case of several different annotators performing annotation of a large-sized corpus. Evaluation of annotation reliability traditionally relies on calculating inter-annotator agreement coefficients. The improvement of annotations consistency in the presented work relies on an automatic method for standardizing argumentative annotations in form of argumentation graphs with two node types: information nodes with the text content of argumentative statements (premises, conclusions) and schemes nodes indicating reasoning models (from Walton’s classification) at the base of connections between statements. The modification of annotations consists in, first, removal of leaf information nodes present only in one annotation version of a text, and second, in replacement of disagreement-causing schemes in accordance with a hierarchy of pairwise rules based on functional properties of these schemes. Calculation of Krippendorff’s alpha shows a considerable increase of inter-annotator agreement after the automatic modification of the corpus containing 100 argumentative annotations for 50 scientific articles. Classification scores between the unmodified and modified corpus exhibit a limited 2–5% increase of F-measure for identifying three analyzed schemes (VerbalClassification, PartToWhole, CorrelationToCause). We conclude that, for evaluating efficiency of automatic annotations unification methods, the improvement of inter-annotator agreement coefficients acts as a less reliable measure than the dynamic of experimental classification scores between the unmodified and modified datasets.

Read the paper · More papers on PaperTik