Machine Learning of Syntactic Attachment from Morphosyntactic and Semantic Co-occurrence Statistics
Szymon AcedaÅ„ski, Adam Slaski, Adam Przepiórkowski · Meeting of the Association for Computational Linguistics · 2012
The paper presents a novel approach to extracting dependency information in morphologically rich languages using co-occurrence statistics based not only on lexical forms (as in previously described collocation-based methods), but also on morphosyntactic and wordnet-derived semantic properties of words. Statistics generated from a corpus annotated only at the morphosyntactic level are used as features in a Machine Learning classifier which is able to detect which heads of groups found by a shallow parser are likely to be connected by an edge in the complete parse tree. The approach reaches the precision of 89% and the recall of 65%, with an extra 6% recall, if only words present in the wordnet are considered.