Improving Named Entity Translation Combining Phonetic and Semantic Similarities
Fei Huang, Stephan Vogel, Alex Waibel · KITopen · 2004
This paper describes an approach to translate rarely occurring named entities (NE) by combining phonetic and semantic similarities. The phonetic similarity is estimated from a surface string transliteration model, and the semantic similarity is calculated from a context vector semantic model. Given a source (Chinese) NE and its context, this approach first generates queries in the target (English) language according to the context translation hypotheses, then searches for relevant documents from a target language corpus. Target NEs in retrieved documents are compared with the source NE based on their phonetic and contextual semantic similarities, and the bestmatched one is selected as the correct translation. Experiments show that this approach achieves 67% accuracy on translating rarely pccuring NEs, and consistently improves the translation quality on different tasks over a state-of-the-art statistical machine translation system.