Whitepaper of NEWS 2010 Shared Task on Transliteration Generation

Haizhou Li, Anju Manakkakudy Kumaran, Min Zhang, Vladimir Pervouchine · 2010

Transliteration is defined as phonetic translation of names across languages. Transliteration of Named Entities (NEs) is necessary in many applications, such as machine translation, corpus alignment, cross-language IR, information extraction and automatic lexicon acquisition. All such systems call for high-performance transliteration, which is the focus of shared task in the NEWS 2010 workshop. The objective of the shared task is to pro-mote machine transliteration research by providing a common benchmarking plat-form for the community to evaluate the state-of-the-art technologies. 1 Task Description The task is to develop machine transliteration sys-tem in one or more of the specified language pairs being considered for the task. Each language pair consists of a source and a target language. The training and development data sets released for each language pair are to be used for developing a transliteration system in whatever way that the participants find appropriate. At the evaluation time, a test set of source names only would be released, on which the participants are expected to produce a ranked list of transliteration candi-dates in another language (i.e. n-best translitera-tions), and this will be evaluated using common metrics. For every language pair the participants must submit at least one run that uses only the data provided by the NEWS workshop organisers in a given language pair (designated as “standard” run, primary submission). Users may submit more “stanrard ” runs. They may also submit several “non-standard ” runs for each language pair that

Read the paper · More papers on PaperTik