Cross-Domain Dutch Coreference Resolution

Orphée De Clercq, Véronique Hoste, Iris Hendrickx · Ghent University Academic Bibliography (Ghent University) · 2011

This article explores the portability of a coreference resolver across a variety of eight text genres.Besides newspaper text, we also include administrative texts, autocues, texts used for external communication, instructive texts, wikipedia texts, medical texts and unedited new media texts.Three sets of experiments were conducted.First, we investigated each text genre individually, and studied the effect of larger training set sizes and including genre-specific training material.Then, we explored the predictive power of each genre for the other genres conducting cross-domain experiments.In a final step, we investigated whether excluding genres with less predictive power increases overall performance.For all experiments we use an existing Dutch mention-pair resolver and report on our experimental results using four metrics: MUC, B-cubed, CEAF and BLANC.We show that resolving out-of-domain genres works best when enough training data is included.This effect is further intensified by including a small amount of genre-specific text.As far as the cross-domain performance is concerned we see that especially genres of a very specific nature tend to have less generalization power.

Read the paper · More papers on PaperTik