ARGUS: A Neuro-Symbolic System Integrating GNNs and LLMs for Actionable Feedback on English Argumentative Writing

Lei Yang, Shuo Zhao · Systems · 2025

English argumentative writing is a cornerstone of academic and professional communication, yet it remains a significant challenge for second-language (L2) learners. While Large Language Models (LLMs) show promise as components in automated feedback systems, their responses are often generic and lack the structural insight necessary for meaningful improvement. Existing Automated Essay Scoring (AES) systems, conversely, typically provide holistic scores without the kind of actionable, fine-grained advice that can guide concrete revisions. To bridge this systemic gap, we introduce ARGUS (Argument Understanding and Structured-feedback), a novel neuro-symbolic system that synergizes the semantic understanding of LLMs with the structured reasoning of Graph Neural Networks (GNNs). The ARGUS system architecture comprises three integrated modules: (1) an LLM-based parser transforms an essay into a structured argument graph; (2) a Relational Graph Convolutional Network (R-GCN) analyzes this symbolic structure to identify specific logical and structural flaws; and (3) this flaw analysis directly guides a conditional LLM to generate feedback that is not only contextually relevant but also pinpoints precise weaknesses in the student’s reasoning. We evaluate ARGUS on the Argument Annotated Essays corpus and on an additional set of 150 L2 persuasive essays collected from the same population to augment training of both the parser and the structural flaw detector. Our argument parsing module achieves a component identification F1-score of 90.4% and a relation identification F1-score of 86.1%. The R-GCN-based structural flaw detector attains a macro-averaged F1-score of 0.83 across the seven flaw categories, indicating that the enriched training data substantially improves its generalization. Most importantly, in a human evaluation study, feedback generated by the ARGUS system was rated as consistently and significantly more specific, accurate, actionable, and helpful than that from strong baselines, including a fine-tuned LLM and a zero-shot GPT-4. Our work demonstrates a robust systems engineering approach, grounding LLM-based feedback in GNN-driven structural analysis to create an intelligent teaching system that provides targeted, pedagogically valuable guidance for L2 student writers engaging with persuasive essays.

Read the paper · More papers on PaperTik