Named Entity Transliteration and Discovery in Multilingual Corpora

Alexandre Klementiev, Dan Roth · The MIT Press eBooks · 2008

This chapter presents a novel algorithm for cross-lingual multiword name entity (NE) discovery in a bilingual weakly temporally aligned corpus. It shows that using two independent sources of information (transliteration and temporal similarity) together to guide NE extraction yields better performance than using them alone. The algorithm requires almost no supervision or linguistic knowledge. The algorithm was evaluated on an English-Russian corpus, and showed a high level of NE discovery in Russian.

Read the paper · More papers on PaperTik