An Effective and Robust Framework for Transliteration Exploration

Ea-Ee Jan, Niyu Ge, Shih-Hsiang Lin, Berlin Chen · 2011

Transliteration is the process of proper name translation based on pronunciation. It is an important process in many multilin-gual natural language tasks. A common and essential component of transliteration approaches is a verification mechanism that tests if the two names in different lan-guages are translations of each other. Al-though many transliteration systems have verification as a component, verification as a stand-alone problem is relatively new. In this paper, we propose a simple, effective and robust training framework for the task of verification. We show the many appli-cations of the verification techniques. Our proposed method can operate on both pho-nemic and orthographic inputs. Our best re-sults show that a simple, straightforward orthographic representation is sufficient and no complex training method is needed. It is effective because it achieves remark-able accuracies. It is robust because it is language-independent. We show that on Chinese and Korean our technique achieves equal error rate well below 1 % and around 1 % for Japanese using 2009 and 2010 NEWS transliteration generation share task dataset. Our results also show that the or-thographic system outperforms the phone-mic system. This is especially encouraging because the orthographic inputs are easier to generate and secondly, one does not need to resort to more complex training al-gorithm to achieve excellent results. This approach is integrated for proper name based cross lingual information retrieval without translation. 1

Read the paper · More papers on PaperTik