Morphosyntactically-Informed Coreference Resolution for Persian with Adaptive Pruning and Global Context Aggregation

Hassan Haji Mohammadi, Alireza Talebpour, Ahmad Mahmoudi-Aznaveh, Samaneh Yazdani · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

Coreference resolution in Persian, a task critical to natural language understanding, presents unique challenges due to the language's pro-drop tendencies, flexible word order, and rich morphosyntactic agreement system. This study introduces the first end-to-end (e-2-e) neural architecture for comprehensive Persian coreference resolution, encompassing pronominal, nominal, and named entity mentions. The system leverages ParsBERT and innovatively integrates Joint Mention Detection and Type Classification (JMDTC), an Adaptive Antecedent Pruning Threshold (AAPT), Morphosyntactically-Informed Attention (MIA), and Cross-Segment Coreference with Global Context Aggregation (CS-GCA). By jointly optimizing mention detection and antecedent linking, the system surpasses traditional pipelined approaches, eliminating the need for handcrafted features and complex syntactic parsers. A CoNLL average F1-score of 76.16% was achieved by the system on the Mehr corpus, which represents a 4.03-point improvement compared with the previous state-of-the-art. Furthermore, it demonstrates robust generalization, achieving a CoNLL average F1-score of 74.20% on the RCDAT corpus (evaluated using the Uppsala test set). These findings facilitate scalable coreference resolution in low-resource languages presenting similar morphosyntactic challenges.

Read the paper · More papers on PaperTik