Vietnamese Text Accent Restoration with Statistical Machine Translation

Luan-Nghia Pham, Viet-Hong Tran, Vinh-Van Nguyen · Institutional Repositories DataBase (IRDB) · 2013

Vietnamese accentless texts exist on parallel with official vietnamese documents and play an important role in instant message, mobile SMS and online searching.Understanding correctly these texts is not simple because of the lexical ambiguity caused by the diversity in adding diacritics to a given accentless sequence.There have been some methods for solving the vietnamese accentless texts problem known as accent prediction and they have obtained promising results.Those methods are usually based on distance matching, n-gram, dictionary of words and phrases and heuristic techniques.In this paper, we propose a new method solving the accent prediction.Our method combine the strength of previous methods (combining n-gram method and phrase dictionary in general).This method considers the accent predicting as statistical machine translation (SMT) problem with source language as accentless texts and target language as accent texts, respectively.We also improve quality of accent predicting by applying some techniques such as adding dictionary, changing order of language model and tuning.The achieved result and the ability to enhance proposed system are obviously promising.

Read the paper · More papers on PaperTik