Collaborative entity extraction and translation

Heng Ji, Ralph Grishman · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2009

Entity extraction is the task of identifying names and nominal phrases (‘mentions’) in a text and linking coreferring mentions. We propose the use of a new source of data for improving entity extraction: the information gleaned from large bitexts and captured by a statistical, phrase-based machine translation system. We translate the individual mentions and test properties of the translated mentions, as well as comparing the translations of coreferring mentions. The results provide feedback to improve source language entity extraction. Experiments on Chinese and English show that this approach can significantly improve Chinese entity extraction (2.2%relative improvement in name tagging F-measure, representing a 15.0 % error reduction), as well as Chinese to English entity translation (9.1 % relative improvement in F-measure), over state-of-the-art entity extraction and machine translation systems.

Read the paper · More papers on PaperTik