MSCAW-coref: Multilingual, Singleton and Conjunction-Aware Word-Level Coreference Resolution

Houjun Liu, John T. Bauer, Karel D’Oosterlinck, Christopher E. Potts, Christopher D. Manning · 2024

Modern multi-lingual coreference resolution approaches largely focus on the clustering of mention spans, leading to quartic complexity in the choice of both spans and span links.The recently published CAW-coref reduces coreference complexity to quadratic while still attaining 97.9% of SOTA performance through a word-level approach on the English OntoNotes slice.Naively extending the CAW-coref algorithm towards multiple languages on the CorefUD dataset results in a lackluster 77.4% of SOTA performance.We find this is due to annotation differences across OntoNotes and CorefUD-the latter features singletons which CAW-coref is not able to classify.In response, we introduce MSCAW-coref, which extends CAW-coref to work in a multilingual setting and accounts for singleton mentions.We demonstrate that MSCAW-coref attains 95.7% of SOTA performance on CorefUD while being substantially more efficient.Our algorithmic contribution towards accounting for singletons is a major driver of performance.Finally, we discuss the cross-linguistic generalization capability of our approach.We release the models, code, and a package for performing coreference analysis for the community as a part of Stanza (https://github.com/ stanfordnlp/stanza).

Read the paper · More papers on PaperTik